Skewcy's Blog


jev.chat

A new model called Jev came out recently, and I got curious. Unlike an LLM it returns structured output by construction, the same shape every time, and it answers in a few hundred milliseconds. My day job is real-time systems, and a model with those two properties looks to me like the first one you could actually put inside a control loop. To try it out, I built a chatbot on it that only speaks emoji: https://jev.chat

jev.chat: pizza or sushi, answered in emoji. -3quarterwidth

A default reply costs about $0.00014 and comes back in under half a second. The site runs on Cloudflare Workers, the code is on GitHub [1], and most of it was written together with Claude Code over two days.

1. What Jev Is

Jev is a “decision model”. You send it a state (any JSON) and a set of typed questions about that state, and it returns an answer to each question. There is no prompt and no completion. The two question types I use:

A request looks like this:

{
  "state": { "message": "My package arrived and the mug inside is broken" },
  "questions": {
    "department": { "type": "choice", "instructions": "Which team should handle this message?",
                    "criteria": { "refunds": null, "shipping": null, "sales": null } },
    "urgency":    { "type": "score",  "instructions": "How urgent is it?",
                    "criteria": ["low", "medium", "high"] }
  }
}

And this is the answer, 353 ms later:

{
  "department": { "probabilities": { "refunds": 0.52, "shipping": 0.48, "sales": 0 } },
  "urgency":    { "probabilities": { "low": 0.17, "medium": 0.80, "high": 0.03 } }
}

That is the entire interface. The output is always the probabilities over the options you gave, so Jev can never produce text, and an emoji-only bot is a natural fit: the output alphabet is whatever you put in the options.

Jev is served through OpenRouter’s System One endpoint [2]. The idea of chatting with it in emoji comes from a small Github project [3]; But i feel talk with limited emojis maybe a better suit for Jev compared to human language

2. Two Ways to Speak

2.1 Multi-shot mode

The first version was the obvious one. Keep asking Jev “what is the next emoji?” until the reply is done:

question ──> Jev: next emoji? ──> 🐉
             Jev: next emoji? ──> 🐉🦸
             Jev: next emoji? ──> 🐉🦸⚔️
             Jev: next emoji? ──> 🐉🦸⚔️🔥
             ...

One detail made this work. The options are not single emoji. Each option is the reply so far plus one candidate, shown as a string:

"🐉🦸⚔️"   "🐉🦸🔥"   "🐉🦸🏰"   "🐉🦸🧙"   ...

This works much better because it builds the relationshp of emojis. A bare 🔥 means nothing on its own; 🐉🦸🔥 is a candidate reply that Jev can judge as a whole, against the question and against the other candidates.

2.2 Single-shot mode

However, multi-shot costs one round trip per emoji, and every round trip carries the question, the conversation and every candidate: an 11-emoji reply was 14 requests and about 10k tokens. It also has to stream, so the user watches the reply arrive one emoji at a time. So I built a second mode.

Ask once. “Which emoji belong in the reply?” over all the candidates gives a distribution, and the distribution is the reply: take the most probable emoji, as many as the reply should be long. How long is a second question in the same request, a Score over four length bands (1–3, 3–6, 6–12, 12–24 emoji). Jev picks the band, the exact count is drawn inside it. Two questions, one request, a fifth of the tokens, no streaming.

Top, multi-shot: one question per emoji, every option is the reply so far plus one candidate. Bottom, single-shot: one question over the candidates, the distribution is read off in probability order to the chosen length. -3quarterwidth

To save tokens, single-shot is the default. Multi-shot reads better, so it stayed as the Think toggle in the composer, with its own daily allowance.

3. Challenges

Single-shot mode sounds too simple to work, and at first it did not. I actually spent most of the time to fix single-shot mode with several versions, the eventual design looks like this:

A message goes through request 1 (which groups, how long) and request 2 (which emoji in each group); the default mode then asks request 3 where each emoji belongs and sorts, think mode asks one request per emoji. -3quarterwidth

3.1 One question over all the candidates does not scale

The first single-shot version asked one question over a candidate set of about 170 emoji. Two problems showed up immediately.

  1. Wasted tokens. Jev has no memory between calls, so every request carries all the candidates, whether the message is about dragons or dinner. The candidates were most of the request.
  2. Bias. Only the top 25 or so options of a Choice ever receive any mass. With 170 options, most of the candidates can never be chosen, and the winners are always the same popular emoji. Growing the candidate set to the full emoji set would make both problems worse, not better.

So the candidates are asked in two tiers. Request 1 asks which groups of emoji the reply draws from (32 groups, 932 emoji in total). Request 2 asks, for each group that earned a slot, which emoji inside that group belong in the reply. A group question is 15 to 45 options, and a message only opens the groups it needs: a story about a knight touches fantasy, actions and reptiles_sea and never sees the 400 emoji about food and weather.

Request 1 puts a share on each of 32 groups; only the groups that earn a slot are asked about in request 2, each over its own emoji. -3quarterwidth

The groups come from Unicode. emoji-test.txt [4] sorts every emoji into subgroups (face-smiling, animal-mammal, food-fruit, …), and I merged those into 32 groups that are mutually exclusive: no emoji appears in two of them. This matters more than it sounds. If 🐉 lived in both fantasy and reptiles_sea, Jev would split its probability between the two groups, and both would look weaker than they are.

3.2 Splitting the reply between groups

Request 1 returns a share for every group. Request 2 needs to know how many emoji to take from each. Formally: the reply has $L$ slots (say $L = 24$ for a story), group $g$ has share $p_g$ with $\sum_g p_g = 1$, and we want integer counts $n_g$ with $\sum_g n_g = L$ that follow the shares as closely as possible.

This is the problem a parliament has after an election, and I used the method many of them use, Sainte-Laguë [5]. Start with every $n_g = 0$ and hand out the $L$ slots one at a time; each slot goes to the group with the largest

\[\frac{p_g}{2 n_g + 1}.\]

A group’s claim shrinks every time it receives a slot, so a group with 2% earns one slot in a 24-emoji story and nothing in a 3-emoji reaction. There is no threshold, no minimum and no quota, and the same rule serves both cases. The same method fills the slots inside a group from request 2’s distribution: highest probability first, no emoji twice.

The peaked-distribution problem. A Choice is a “which one” question, and Jev answers it that way: it piles almost all the mass on the single best option even when the honest answer is “several of these”. For the knight story it says:

Group $p_g$
fantasy 0.91
actions 0.07
reptiles_sea 0.02

Apportion 24 slots on that and you get 22 heroes and wizards, 2 actions, and nothing for the group that holds the dragon:

Group $p_g$ $n_g$ (raw) $n_g$ (after square root)
fantasy 0.91 22 17
actions 0.07 2 5
reptiles_sea 0.02 0 2

The raw allocation produced 🦸👹🧙☠️👾⚔️. A knight story with no dragon.

The fix is to flatten the shares before apportioning:

\[\tilde p_g = \frac{\sqrt{p_g}}{\sum_h \sqrt{p_h}}.\]

Small shares grow, the big one shrinks, the ranking stays. With it the same story becomes 🦸👹🧙⚔️🔥🛡️🐉🐲, and the fox story gets its ending: 🦊🏖️🏝️🐳🐋🌙📖🛌😊 instead of stopping at the whale.

Three groups with shares 0.91, 0.07 and 0.02 fill 24 slots. Raw shares give 22, 2 and 0 slots and a story without a dragon; square-rooted shares give 17, 5 and 2 and a story with a fight and a dragon. -3quarterwidth

3.3 Order: ask Jev where each emoji goes

After requests 1 and 2 there is a bag of emoji sorted by probability. Probability is not narrative. 🎉 should open a congratulation and 🏆 should not; 🛌 belongs at the end of a bedtime story, not the middle.

Request 3 asks Jev directly. One Score question per emoji, five levels, first / early / middle / late / last, with the whole set in the state so Jev sees what it is ordering. The reply is sorted by the expected level. For the fox story:

Emoji expected position (0 = first, 4 = last)
🦊 0.09
🏖️ 1.00
🏝️ 1.28
🐳 1.94
🐋 2.03
🐬 2.16
🌙 2.98
📖 3.00
🛌 3.00
😊 3.99

That gives 🦊🏖️🏝️🐳🐋🌙📖🛌😊: the fox, the beach, the sea creatures, night, the book, bed, a smile. Jev cannot write a story, but it knows what order a story goes in. The same request moves 🎉 in front of 🏆 every time (🏆🎉🎊🥳 → 🎉🏆🎊🥳).

Ten emoji placed on a line from first to last at the position Jev expects for each: the fox first, the beach and the sea early, night and the book late, bed near the end, a smile last. -3quarterwidth

Before request 3 existed there was a special “story path” with fixed who / where / what / ending questions, which decided the order by construction. Once Jev could order things itself, the story path was deleted, along with a topic threshold, a scene weight, and two other numbers I had tuned by hand. The rule I ended up with: every decision is an ordinary question to Jev, and the code only counts.

4. Cost, and How Jev Compares

The take away here is Jev is the fastest and cheaper model but it is definiately not a good story teller. Of course this is not what Jev is designed for :)

A default reply is three requests and about 3,000 input tokens, $0.00013 at OpenRouter’s price. Request 1 goes out for every message, so it was compacted from 1.7k to 1.15k tokens: short group keys, a few words per group, one line per length band. A Think reply is one request per emoji, 4k to 11k tokens, so it is capped at ten per visitor per day. The site itself is a Cloudflare Worker with static assets, a Durable Object that counts the day’s messages against a $10 budget, and an answer cache for first messages (the starter prompts repeat all day; a cache hit costs nothing and the reply still varies, because the length is drawn fresh).

The honest question is whether any of this beats a normal language model told to answer in emoji. So I asked three of them the same eight prompts through OpenRouter, with one system prompt: “You are Jev, a friendly chat companion who speaks only in emoji. Reply with emoji only: no words, no letters, no punctuation.” Jev ran through the same code as the site, also straight against OpenRouter, so nothing was cached. The three models are gpt-4o-mini from the cheap tier, claude-haiku-4.5 because it costs about the same per reply as Jev, and claude-sonnet-5 as a frontier reference. Everything below is one run, unedited.

Prompt Jev, default Jev, think gpt-4o-mini claude-haiku-4.5 claude-sonnet-5
Once upon a time a knight rode out to fight a dragon. Tell that story. 🦸🧙👹⚔️🔥🐉 🐉⚔️🦸🔥🤺👹🐲🧙👾 🛡️🐉🏇🔪💥🌌👑✨🗡️🎇🔥🏰🌲👀🤺🎉👸❤️ ⚔️🐉🏰👑 🐴🧝‍♂️➡️🏔️ 🔥💨😨 ⚔️⚔️🗡️💥🔥 😰⚔️🐉💪 🗡️✨💫⚡ 🐉💥💀 😊🧝‍♂️👑🎉 🏆✨🍾 🏰🤴⚔️🐴➡️🌲🌲🌲🏔️🐉 🤴😨🐴🛑 🐉🔥🔥🔥 🤴🛡️➡️🐉 ⚔️💥🐉 🐉😡🔥💨 🤴🏃‍♂️💨🛡️🔥 🐉👀🤴 🤴⚔️➡️🐉👁️ 🐉😵💫 🤴⚔️⚔️⚔️➡️🐉 🐉💀 🤴🏆🎉 🤴🐴🏰➡️ 👑😍🤴💕 🎉🍻🎊👑🤴
A bedtime story: a little fox travels to see the sea for the first time 🦊🏖️🏝️🐳🌊⛵🦸📖🛏️🛌 🦊🏖️📖🐬🐳🛌🏝️🛏️ 🦊🛤️🌳🌌✨🛶➡️🌊🏖️🐚☀️🌊💦😮🦈🐠🏄‍♂️🏄‍♀️💖🥳🌙💤 🦊✨🌙 🚶‍♂️➡️🌊 🏞️🌲🌲 ➡️ 🏖️ 👀🌊😲✨ 🌊🌊💙💎 🦊💕🌅 🐚🦀🐠 ➡️ 👀 🌊🎐💤 🦊😴🌙⭐ 💭🦊💕🌊 🏠➡️❤️ 🦊🌙✨🚶‍♂️🌲🌲🌲🏞️🌾🚶‍♀️🌄🐾🐾🐾🛤️🌳🌳🐦☁️☀️🚶‍♂️⛰️👀🌊😲💙💙💙🌊🌊🐚🐠🐬🌅🦊😊🐾🏖️👣👣🌊🦊💦😄🐟🐡🦀🌊🦊🤣💧🌇🧡🦊😌🏖️🌌⭐⭐⭐🌙🦊😴💤💭🌊💙🌟
Write me a short poem about maple leaves in Kyoto in autumn 🍁🍂⛰️🗻🏔️ 🍁🌅🍂🎨⛰️🖼️🗻🏔️ 🍁🌿🍂🌅🍁🌳🍂🍁🏯💕🍂🍁✨🍃🍁🤗 🍁🏯🍂✨ 🍁💛🌅🍁 🍁🌪️💫🍁 🍁🛤️😊🍁 🍁❤️🍂✨ 🍁🍁🍁🍂🍂🎏🏯🍁🍂🌬️ 🍂🍁🎋🍁🍂🏯🍁🌅🍂🍁 🍁🏯⛩️🍂🍁🍂🌬️🍁🍂🍁 🍂🍁🎐🍁🍂🌙🏯🍁🍂🍁
Our family cat passed away this morning 😢😭🥹🐱 😢🖤🐱 😿💔🌈🐾😢 😢💔😿🤍✨🕊️😭💐 😿💔🐾😢🕊️🌈🐈‍⬛💕🙏
Any tips for falling asleep faster tonight? 🛌🧘🛏️🕯️😴🌙⏰ 💤🛏️🧘💡🌙🕯️🛌 😴🌙✨🛏️🕯️🍵🧘‍♂️📱❌🚫💤💤💤 🛏️😴🌙⏰💤 🧘‍♂️🕯️🌬️ 📵☕❌ 🥗🍌✨ 🎵🎧🧊 😊💚👍 🌙😌🛏️📵📱🚫☕😴🕯️🧘‍♂️💨📖👀⬇️🧦🥤💧🌡️❄️🎵🎧😊✨⭐️🌌😴💤
Guess what, I got the promotion! 🎉🏆🎊🥇🥳🤩 🎉🥳 🎉🥳🙌💼🎈👏🎊✨ 🎉🎊👏🥳🌟💼🚀🎈😊👍 🎉🎊👏🥳🍾🙌🚀⭐🥇👍
What should I pack for a week of camping in the mountains? 🏕️⛺⛰️🎒👕🧰 🏕️🎒🏔️👢🧥⛺👕🧰🛠️⛰️🚶 🎒🏕️⛺🔥🥾🧢🌄🧗‍♂️🍽️🥗💧🍏🌈📸🎣🧳✨ 🎒🏕️⛺🧥🧢 👖👕🧦👟🧤 🧴🧼🪥🧻 🍎🥫🧂💧 🔦🪵⚒️🔪 🧭📱🔋 😴💤🛏️ 🎣🚶‍♂️🧗 ☀️🌧️❄️ 😊✨ ⛺🎒🔥🪵🔦🧭🗺️👖🧥🧤🧢🕶️🥾🧦👕🩳🌂☔❄️🌡️🔋🥤💧🍫🥜🍳🥓🍞🔪🍽️🥄🩹💊🧻🧼🪥🚿🐻🦟🕸️🌲⛰️🌙⭐📱🔌📷🧦🧤
My laptop crashed and I lost my whole presentation 😭😢😥😟💻 😢💻😥🤦 😱💻💔😩🖥️😕 😰😩💔🖥️⚠️😟 😤🤔💭🔧🆘 ☁️💾📱🤷‍♂️❓ 🤗💪🆙😌✨ 😱💻💥😢🙏🔍📁☁️💾🔄😔🤗
Model cost per reply time per reply Jev calls emoji per reply
Jev, default $0.00014 0.4 s 3 6.1
Jev, think $0.00024 1.1 s 8 6.5
gpt-4o-mini $0.00003 1.6 s 1 13.6
claude-haiku-4.5 $0.00045 1.8 s 1 24.5
claude-sonnet-5 $0.00135 3.5 s 1 36.3

Cost and time are averages over the eight prompts, measured at the client; every reply from every model was emoji only.

Reference

[1] jev.chat source: https://github.com/ChuanyuXue/jev.chat

[2] OpenRouter, System One endpoint (TypeSafe Jev): https://openrouter.ai/typesafe/jev-1.13

[3] Kyle Pena, jevchat: https://github.com/kyle-pena-nlp/jevchat

[4] Unicode emoji-test.txt: https://unicode.org/Public/emoji/latest/emoji-test.txt

[5] Sainte-Laguë method: https://en.wikipedia.org/wiki/Sainte-Lagu%C3%AB_method

[6] OpenAI, Introducing Structured Outputs in the API (2024): https://openai.com/index/introducing-structured-outputs-in-the-api/

[7] Instruction-Following Evaluation in Function Calling for Large Language Models (2025), Tables 3 and 4: https://arxiv.org/abs/2509.18420


skewcy@gmail.com