Most of my coding agent work now runs through more than one agent. A manager agent brainstorms with me, starts an executor, checks on it, and later brings in a reviewer. Sometimes I resume a session the next morning. Sometimes the manager is running on a VPS and I talk to it from Telegram.
Every one of those moves is a handoff. Each time, the same question comes up: what does the next agent actually need to know?
The obvious answer is the transcript. That doesn’t work for long. This post is about a small AI DevKit command that tries to answer the question better, and about Jev, the model behind it.
The problem is state handoff, not context length
Coding agent sessions get long quickly. An hour of implementation work can produce hundreds of messages: the original instruction, clarifications, file reads, test runs, failed attempts, retries, status updates, and the occasional “let me check that again”.
Somewhere in there are the facts that matter:
- what the user asked for, including the constraints they added halfway through,
- which decisions were made, and why,
- which files changed,
- which commands ran and what they proved,
- what is still blocked,
- what the next step is.
They are mixed with a lot of noise. The same npm test output appears four times. A path the agent tried and abandoned takes up twenty messages. Status chatter (“Now I’ll look at the service layer”) says nothing a later reader needs.
When I work with a single agent, this is mostly an annoyance. The harness compacts the context, some detail gets lost, and I repeat myself.
In orchestration it turns into a reliability problem. If the manager hands an executor’s state to a reviewer and the handoff drops a constraint, the reviewer approves the wrong thing. If a resumed session loses the validation evidence, the agent reruns work or, worse, claims something passed when nobody checked. A bigger context window doesn’t fix this. What fixes it is choosing the right state to pass on.
Jev in one sentence
Jev is a model from TypeSafe that takes unstructured input and returns typed decisions, fast.
TypeSafe calls it a System One Model and sums up the idea as “unstructured state in, typed probabilistic decisions out”. You don’t ask Jev to write. You give it a piece of state and a set of typed questions: pick one of these categories, give me a probability for this yes/no. It answers in that shape, with a probability or confidence attached.
According to TypeSafe, that narrow focus buys a few things:
- Type-safe answers. The output always matches the schema you asked for. TypeSafe goes further and says Jev “can’t hallucinate” for these queries, because the answer is constrained to the options you defined.
- Calibrated confidence. Every answer comes with a probability, and TypeSafe says higher confidence means higher accuracy.
- Latency. TypeSafe reports 70–500 ms end to end, and says that is 40 to 200x faster than frontier chat LLMs for “System One shaped” queries.
I haven’t benchmarked these numbers carefully, so treat them as TypeSafe’s claims. What matters for this post is the shape: Jev isn’t competing with chat LLMs at writing code or explaining things. It does the small, frequent judgement calls that sit between those bigger steps.
Structured output isn’t new
I want to be careful here, because it is easy to present this as a brand-new idea. It isn’t.
Jev doesn’t introduce the idea of structured AI output. Engineers have been moving toward “AI as a structured-data engine” for a while, with JSON mode, tool calling, and schema-constrained decoding.
Last year I wrote The turning point in AI. The first point in that post was that AI isn’t new. What changed was access, and with it, how engineers build.
The example I used was a weather app. The old way: an engineer writes the integration by hand. The early-AI way: the engineer asks an LLM to help write that same integration code. The shift I cared about was a third option: ask the model directly for the temperature, humidity, wind speed, and chance of rain for a location, in a structured format, and build the app around that output. In that post I put it as “Treat AI as an engine you can build around. Prompting becomes programming”.
Structured output was the important part then, and it still is. Once a model returns clean JSON instead of a paragraph, you can put it inside a program. You can branch on it, store it, test it, and pass it to the next step.
What Jev changes is the cost of each call and what the call returns. If a typed decision takes a few hundred milliseconds and comes back with a calibrated probability, you can afford to ask it many times inside a loop, not just once at the edge of a user interaction. Classifying every message in a session stops being an expensive batch job and becomes something you can simply do.
That’s the part I find exciting: an old idea that is now cheap enough to use in places where it used to be too slow.
Why session compact needs a typed decision layer
A normal summary keeps the story of a session. For continuing work, I don’t need the story. I need the operational facts.
So the first design question for session compact was: what should survive?
Keep:
- user instructions and constraints,
- decisions,
- code changes,
- command evidence,
- validation evidence,
- blockers and open questions,
- the next step,
- facts worth proposing as long-term memory.
Drop:
- routine status updates,
- duplicated tool output,
- abandoned exploration that led nowhere,
- anything sensitive, such as tokens, keys, and credentials.
Written that way, this is a classification problem, not a writing problem. For every event in the session I want answers to four questions:
- Which category is this?
- How important is it for continuing the work?
- Should it survive compaction?
- Does it contain something that shouldn’t be stored?
Each question has a small, fixed set of possible answers. That is exactly what Jev is built for. The categories AI DevKit uses are:
- user_instruction
- decision
- code_change
- command_evidence
- validation_evidence
- blocker
- next_step
- memory_candidate
- discard
Once every event has typed labels, building the compact artifact doesn’t need another model call. Plain code can do it.
How AI DevKit implements it
The command lives under the existing agent session namespace:
# Find the session you want to compactnpx ai-devkit@latest agent sessions --all# Human-readable handoffnpx ai-devkit@latest agent session compact --id <session-id> --format markdown# Machine-readable handoff for another agent or a scriptnpx ai-devkit@latest agent session compact --id <session-id> --format json
Markdown is the default. If the same session ID exists for more than one provider, --type narrows the lookup (claude, codex, gemini_cli, opencode, pi, and others).
Here is the pipeline:

A real run: compacting the session that built this feature
For a real example, I compacted the Codex session that built this feature. A manager agent handed the requirement to a Codex executor, and that executor took it through requirements, design, planning, implementation, tests, commit, rebase, and send PR. It ran for a few hours of wall-clock time.
The adapter returned 55 messages (9 user, 40 assistant, 6 system). Jev classified all of them sequentially in about 0.36s.
Here is how the size changed at each stage. Token counts use the o200k_base tokenizer, so treat them as estimates rather than any specific model’s exact count.
| Stage | File size | Tokens |
|---|---|---|
Raw Codex rollout file (.jsonl) | 5.6 MB | ~1.72M |
| Conversation the adapter gave Jev (55 messages) | 96 KB | ~21.6K |
| Compact, Markdown | 27 KB | ~5.9K |
| Compact, JSON | 32 KB | ~7.1K |
The raw file isn’t a fair baseline. Most of its 5.6 MB is event logs, tool calls and outputs, encrypted reasoning, and Codex’s own compaction records, and no model ever reads all of it at once. The numbers I care about are these two:
- Against the conversation Jev saw: 21.6K → 5.9K tokens, about 73% smaller.
- Against what Codex was actually holding in context at the end: 130.6K → 5.9K tokens, about 95% smaller. A fresh agent can start from roughly 6K tokens instead of inheriting a context that was half full of a 258K window.
Why this matters for orchestration
Once agents start working together, communication is part of the product. Here is where compact fits in the loop I use every day:

A few things change once handoffs carry compact state instead of transcripts:
- The manager stays small. It doesn’t have to read every executor transcript to know where things stand. It reads intent, current state, evidence, and next step.
- Reviewers start from facts. They get the constraints and the validation evidence up front, instead of digging through the transcript for them.
- Resuming is cheaper. A stale session can restart from its resume prompt instead of rereading hours of noise.
- Memory gets better input. Memory candidates come out as a separate list, so a memory skill can review them instead of me remembering to extract them.
A good handoff isn’t a longer summary. It’s the right state, chosen carefully. Session compact is a small primitive. Multi-agent setups become dependable through small primitives like this one.
Try it
If you’re building coding-agent workflows, and especially manager/executor/reviewer loops or long-running tasks that span sessions, try AI DevKit.
# Set up AI DevKitnpm i -g ai-devkitai-devkit setup# See the agents and sessions it can discoverai-devkit agent sessions --all# Get a Jev key from TypeSafe, then compact a sessionexport TYPESAFE_API_KEY=YOUR_API_KEY_HEREai-devkit agent session compact --id <session-id>
Run it on a long session you already have and compare the result with what you’d have written by hand for the next agent. If the compact misses something you needed, or keeps something you didn’t, I’d like to hear about it. That feedback is how the categories and filter will improve.
The code is on GitHub. If my sharing is helpful to you, subscribe to my blog. I share what I learn while building real systems with AI in the loop. You can also follow me on X or Threads for more thoughts and ongoing experiments.
Sources
- TypeSafe, Introducing System One Models and Jev. Latency, hallucination, and calibration claims in this post come from this source.
- Codeaholicguy, The turning point in AI, May 2025.
Discover more from Codeaholicguy
Subscribe to get the latest posts sent to your email.