A new model called Jev came out recently, and I got curious. Unlike an LLM it returns structured output by construction, the same shape every time, and it answers in a few hundred milliseconds. My day job is real-time systems, and a model with those two properties looks to me like the first one you could actually put inside a control loop. To try it out, I built a chatbot on it that only speaks emoji: https://jev.chat
A default reply costs about $0.00014 and comes back in under half a second. The site runs on Cloudflare Workers, the code is on GitHub [1], and most of it was written together with Claude Code over two days.
Jev is a “decision model”. You send it a state (any JSON) and a set of typed questions about that state, and it returns an answer to each question. There is no prompt and no completion. The two question types I use:
A request looks like this:
{
"state": { "message": "My package arrived and the mug inside is broken" },
"questions": {
"department": { "type": "choice", "instructions": "Which team should handle this message?",
"criteria": { "refunds": null, "shipping": null, "sales": null } },
"urgency": { "type": "score", "instructions": "How urgent is it?",
"criteria": ["low", "medium", "high"] }
}
}
And this is the answer, 353 ms later:
{
"department": { "probabilities": { "refunds": 0.52, "shipping": 0.48, "sales": 0 } },
"urgency": { "probabilities": { "low": 0.17, "medium": 0.80, "high": 0.03 } }
}
That is the entire interface. The output is always the probabilities over the options you gave, so Jev can never produce text, and an emoji-only bot is a natural fit: the output alphabet is whatever you put in the options.
Jev is served through OpenRouter’s System One endpoint [2]. The idea of chatting with it in emoji comes from a small Github project [3]; But i feel talk with limited emojis maybe a better suit for Jev compared to human language
The first version was the obvious one. Keep asking Jev “what is the next emoji?” until the reply is done:
question ──> Jev: next emoji? ──> 🐉
Jev: next emoji? ──> 🐉🦸
Jev: next emoji? ──> 🐉🦸⚔️
Jev: next emoji? ──> 🐉🦸⚔️🔥
...
One detail made this work. The options are not single emoji. Each option is the reply so far plus one candidate, shown as a string:
"🐉🦸⚔️" "🐉🦸🔥" "🐉🦸🏰" "🐉🦸🧙" ...
This works much better because it builds the relationshp of emojis. A bare 🔥 means nothing on its own; 🐉🦸🔥 is a candidate reply that Jev can judge as a whole, against the question and against the other candidates.
However, multi-shot costs one round trip per emoji, and every round trip carries the question, the conversation and every candidate: an 11-emoji reply was 14 requests and about 10k tokens. It also has to stream, so the user watches the reply arrive one emoji at a time. So I built a second mode.
Ask once. “Which emoji belong in the reply?” over all the candidates gives a distribution, and the distribution is the reply: take the most probable emoji, as many as the reply should be long. How long is a second question in the same request, a Score over four length bands (1–3, 3–6, 6–12, 12–24 emoji). Jev picks the band, the exact count is drawn inside it. Two questions, one request, a fifth of the tokens, no streaming.
To save tokens, single-shot is the default. Multi-shot reads better, so it stayed as the Think toggle in the composer, with its own daily allowance.
Single-shot mode sounds too simple to work, and at first it did not. I actually spent most of the time to fix single-shot mode with several versions, the eventual design looks like this:
The first single-shot version asked one question over a candidate set of about 170 emoji. Two problems showed up immediately.
So the candidates are asked in two tiers. Request 1 asks which groups of emoji the reply draws from (32 groups, 932 emoji in total). Request 2 asks, for each group that earned a slot, which emoji inside that group belong in the reply. A group question is 15 to 45 options, and a message only opens the groups it needs: a story about a knight touches fantasy, actions and reptiles_sea and never sees the 400 emoji about food and weather.
The groups come from Unicode. emoji-test.txt [4] sorts every emoji into subgroups (face-smiling, animal-mammal, food-fruit, …), and I merged those into 32 groups that are mutually exclusive: no emoji appears in two of them. This matters more than it sounds. If 🐉 lived in both fantasy and reptiles_sea, Jev would split its probability between the two groups, and both would look weaker than they are.
Request 1 returns a share for every group. Request 2 needs to know how many emoji to take from each. Formally: the reply has $L$ slots (say $L = 24$ for a story), group $g$ has share $p_g$ with $\sum_g p_g = 1$, and we want integer counts $n_g$ with $\sum_g n_g = L$ that follow the shares as closely as possible.
This is the problem a parliament has after an election, and I used the method many of them use, Sainte-Laguë [5]. Start with every $n_g = 0$ and hand out the $L$ slots one at a time; each slot goes to the group with the largest
\[\frac{p_g}{2 n_g + 1}.\]A group’s claim shrinks every time it receives a slot, so a group with 2% earns one slot in a 24-emoji story and nothing in a 3-emoji reaction. There is no threshold, no minimum and no quota, and the same rule serves both cases. The same method fills the slots inside a group from request 2’s distribution: highest probability first, no emoji twice.
The peaked-distribution problem. A Choice is a “which one” question, and Jev answers it that way: it piles almost all the mass on the single best option even when the honest answer is “several of these”. For the knight story it says:
| Group | $p_g$ |
|---|---|
| fantasy | 0.91 |
| actions | 0.07 |
| reptiles_sea | 0.02 |
Apportion 24 slots on that and you get 22 heroes and wizards, 2 actions, and nothing for the group that holds the dragon:
| Group | $p_g$ | $n_g$ (raw) | $n_g$ (after square root) |
|---|---|---|---|
| fantasy | 0.91 | 22 | 17 |
| actions | 0.07 | 2 | 5 |
| reptiles_sea | 0.02 | 0 | 2 |
The raw allocation produced 🦸👹🧙☠️👾⚔️. A knight story with no dragon.
The fix is to flatten the shares before apportioning:
\[\tilde p_g = \frac{\sqrt{p_g}}{\sum_h \sqrt{p_h}}.\]Small shares grow, the big one shrinks, the ranking stays. With it the same story becomes 🦸👹🧙⚔️🔥🛡️🐉🐲, and the fox story gets its ending: 🦊🏖️🏝️🐳🐋🌙📖🛌😊 instead of stopping at the whale.
After requests 1 and 2 there is a bag of emoji sorted by probability. Probability is not narrative. 🎉 should open a congratulation and 🏆 should not; 🛌 belongs at the end of a bedtime story, not the middle.
Request 3 asks Jev directly. One Score question per emoji, five levels, first / early / middle / late / last, with the whole set in the state so Jev sees what it is ordering. The reply is sorted by the expected level. For the fox story:
| Emoji | expected position (0 = first, 4 = last) |
|---|---|
| 🦊 | 0.09 |
| 🏖️ | 1.00 |
| 🏝️ | 1.28 |
| 🐳 | 1.94 |
| 🐋 | 2.03 |
| 🐬 | 2.16 |
| 🌙 | 2.98 |
| 📖 | 3.00 |
| 🛌 | 3.00 |
| 😊 | 3.99 |
That gives 🦊🏖️🏝️🐳🐋🌙📖🛌😊: the fox, the beach, the sea creatures, night, the book, bed, a smile. Jev cannot write a story, but it knows what order a story goes in. The same request moves 🎉 in front of 🏆 every time (🏆🎉🎊🥳 → 🎉🏆🎊🥳).
Before request 3 existed there was a special “story path” with fixed who / where / what / ending questions, which decided the order by construction. Once Jev could order things itself, the story path was deleted, along with a topic threshold, a scene weight, and two other numbers I had tuned by hand. The rule I ended up with: every decision is an ordinary question to Jev, and the code only counts.
The take away here is Jev is the fastest and cheaper model but it is definiately not a good story teller. Of course this is not what Jev is designed for :)
A default reply is three requests and about 3,000 input tokens, $0.00013 at OpenRouter’s price. Request 1 goes out for every message, so it was compacted from 1.7k to 1.15k tokens: short group keys, a few words per group, one line per length band. A Think reply is one request per emoji, 4k to 11k tokens, so it is capped at ten per visitor per day. The site itself is a Cloudflare Worker with static assets, a Durable Object that counts the day’s messages against a $10 budget, and an answer cache for first messages (the starter prompts repeat all day; a cache hit costs nothing and the reply still varies, because the length is drawn fresh).
The honest question is whether any of this beats a normal language model told to answer in emoji. So I asked three of them the same eight prompts through OpenRouter, with one system prompt: “You are Jev, a friendly chat companion who speaks only in emoji. Reply with emoji only: no words, no letters, no punctuation.” Jev ran through the same code as the site, also straight against OpenRouter, so nothing was cached. The three models are gpt-4o-mini from the cheap tier, claude-haiku-4.5 because it costs about the same per reply as Jev, and claude-sonnet-5 as a frontier reference. Everything below is one run, unedited.
| Prompt | Jev, default | Jev, think | gpt-4o-mini | claude-haiku-4.5 | claude-sonnet-5 |
|---|---|---|---|---|---|
| Once upon a time a knight rode out to fight a dragon. Tell that story. | 🦸🧙👹⚔️🔥🐉 | 🐉⚔️🦸🔥🤺👹🐲🧙👾 | 🛡️🐉🏇🔪💥🌌👑✨🗡️🎇🔥🏰🌲👀🤺🎉👸❤️ | ⚔️🐉🏰👑 🐴🧝♂️➡️🏔️ 🔥💨😨 ⚔️⚔️🗡️💥🔥 😰⚔️🐉💪 🗡️✨💫⚡ 🐉💥💀 😊🧝♂️👑🎉 🏆✨🍾 | 🏰🤴⚔️🐴➡️🌲🌲🌲🏔️🐉 🤴😨🐴🛑 🐉🔥🔥🔥 🤴🛡️➡️🐉 ⚔️💥🐉 🐉😡🔥💨 🤴🏃♂️💨🛡️🔥 🐉👀🤴 🤴⚔️➡️🐉👁️ 🐉😵💫 🤴⚔️⚔️⚔️➡️🐉 🐉💀 🤴🏆🎉 🤴🐴🏰➡️ 👑😍🤴💕 🎉🍻🎊👑🤴 |
| A bedtime story: a little fox travels to see the sea for the first time | 🦊🏖️🏝️🐳🌊⛵🦸📖🛏️🛌 | 🦊🏖️📖🐬🐳🛌🏝️🛏️ | 🦊🛤️🌳🌌✨🛶➡️🌊🏖️🐚☀️🌊💦😮🦈🐠🏄♂️🏄♀️💖🥳🌙💤 | 🦊✨🌙 🚶♂️➡️🌊 🏞️🌲🌲 ➡️ 🏖️ 👀🌊😲✨ 🌊🌊💙💎 🦊💕🌅 🐚🦀🐠 ➡️ 👀 🌊🎐💤 🦊😴🌙⭐ 💭🦊💕🌊 🏠➡️❤️ | 🦊🌙✨🚶♂️🌲🌲🌲🏞️🌾🚶♀️🌄🐾🐾🐾🛤️🌳🌳🐦☁️☀️🚶♂️⛰️👀🌊😲💙💙💙🌊🌊🐚🐠🐬🌅🦊😊🐾🏖️👣👣🌊🦊💦😄🐟🐡🦀🌊🦊🤣💧🌇🧡🦊😌🏖️🌌⭐⭐⭐🌙🦊😴💤💭🌊💙🌟 |
| Write me a short poem about maple leaves in Kyoto in autumn | 🍁🍂⛰️🗻🏔️ | 🍁🌅🍂🎨⛰️🖼️🗻🏔️ | 🍁🌿🍂🌅🍁🌳🍂🍁🏯💕🍂🍁✨🍃🍁🤗 | 🍁🏯🍂✨ 🍁💛🌅🍁 🍁🌪️💫🍁 🍁🛤️😊🍁 🍁❤️🍂✨ | 🍁🍁🍁🍂🍂🎏🏯🍁🍂🌬️ 🍂🍁🎋🍁🍂🏯🍁🌅🍂🍁 🍁🏯⛩️🍂🍁🍂🌬️🍁🍂🍁 🍂🍁🎐🍁🍂🌙🏯🍁🍂🍁 |
| Our family cat passed away this morning | 😢😭🥹🐱 | 😢🖤🐱 | 😿💔🌈🐾😢 | 😢💔😿🤍✨🕊️😭💐 | 😿💔🐾😢🕊️🌈🐈⬛💕🙏 |
| Any tips for falling asleep faster tonight? | 🛌🧘🛏️🕯️😴🌙⏰ | 💤🛏️🧘💡🌙🕯️🛌 | 😴🌙✨🛏️🕯️🍵🧘♂️📱❌🚫💤💤💤 | 🛏️😴🌙⏰💤 🧘♂️🕯️🌬️ 📵☕❌ 🥗🍌✨ 🎵🎧🧊 😊💚👍 | 🌙😌🛏️📵📱🚫☕😴🕯️🧘♂️💨📖👀⬇️🧦🥤💧🌡️❄️🎵🎧😊✨⭐️🌌😴💤 |
| Guess what, I got the promotion! | 🎉🏆🎊🥇🥳🤩 | 🎉🥳 | 🎉🥳🙌💼🎈👏🎊✨ | 🎉🎊👏🥳🌟💼🚀🎈😊👍 | 🎉🎊👏🥳🍾🙌🚀⭐🥇👍 |
| What should I pack for a week of camping in the mountains? | 🏕️⛺⛰️🎒👕🧰 | 🏕️🎒🏔️👢🧥⛺👕🧰🛠️⛰️🚶 | 🎒🏕️⛺🔥🥾🧢🌄🧗♂️🍽️🥗💧🍏🌈📸🎣🧳✨ | 🎒🏕️⛺🧥🧢 👖👕🧦👟🧤 🧴🧼🪥🧻 🍎🥫🧂💧 🔦🪵⚒️🔪 🧭📱🔋 😴💤🛏️ 🎣🚶♂️🧗 ☀️🌧️❄️ 😊✨ | ⛺🎒🔥🪵🔦🧭🗺️👖🧥🧤🧢🕶️🥾🧦👕🩳🌂☔❄️🌡️🔋🥤💧🍫🥜🍳🥓🍞🔪🍽️🥄🩹💊🧻🧼🪥🚿🐻🦟🕸️🌲⛰️🌙⭐📱🔌📷🧦🧤 |
| My laptop crashed and I lost my whole presentation | 😭😢😥😟💻 | 😢💻😥🤦 | 😱💻💔😩🖥️😕 | 😰😩💔🖥️⚠️😟 😤🤔💭🔧🆘 ☁️💾📱🤷♂️❓ 🤗💪🆙😌✨ | 😱💻💥😢🙏🔍📁☁️💾🔄😔🤗 |
| Model | cost per reply | time per reply | Jev calls | emoji per reply |
|---|---|---|---|---|
| Jev, default | $0.00014 | 0.4 s | 3 | 6.1 |
| Jev, think | $0.00024 | 1.1 s | 8 | 6.5 |
| gpt-4o-mini | $0.00003 | 1.6 s | 1 | 13.6 |
| claude-haiku-4.5 | $0.00045 | 1.8 s | 1 | 24.5 |
| claude-sonnet-5 | $0.00135 | 3.5 s | 1 | 36.3 |
Cost and time are averages over the eight prompts, measured at the client; every reply from every model was emoji only.
[1] jev.chat source: https://github.com/ChuanyuXue/jev.chat
[2] OpenRouter, System One endpoint (TypeSafe Jev): https://openrouter.ai/typesafe/jev-1.13
[3] Kyle Pena, jevchat: https://github.com/kyle-pena-nlp/jevchat
[4] Unicode emoji-test.txt: https://unicode.org/Public/emoji/latest/emoji-test.txt
[5] Sainte-Laguë method: https://en.wikipedia.org/wiki/Sainte-Lagu%C3%AB_method
[6] OpenAI, Introducing Structured Outputs in the API (2024): https://openai.com/index/introducing-structured-outputs-in-the-api/
[7] Instruction-Following Evaluation in Function Calling for Large Language Models (2025), Tables 3 and 4: https://arxiv.org/abs/2509.18420
skewcy@gmail.com