Every model improvement creates an evaluation problem. A new checkpoint, a different training recipe, a revised prompt: each produces another candidate that might improve the product. As our training throughput increased, we needed an evaluation process that could keep up.
For our conversational products, we relied on two approaches: Arena-style Elo comparisons and user-level A/B tests. Elo let us compare responses directly. A/B tests let us measure what happened when users actually experienced a model. But there was a gap between the feedback each provided and the iteration loop we wanted.
We built Session AB to fill that gap: an online evaluation system that measures how models perform over sustained interactions with real users. Within a month of building the infrastructure, we had evaluated hundreds of model variants and were testing over a hundred each week. Here’s how we arrived at the design, what running it looks like, and how its results relate to user-level A/B tests.
Why single-response comparisons left a gap
An Elo comparison asks a straightforward question: given the same conversation context, which of two responses does someone prefer? That is useful feedback. But our optimization goals extend beyond the next response. We want models that sustain enjoyable conversations, follow users’ intentions, and develop stories people want to continue.
We found that Elo comparisons favored responses that were more exciting and attention-grabbing. That raised a concern: a response that stands out in a side-by-side comparison may not lead to a better experience as the conversation unfolds.
There is also an off-policy problem: the conversation histories used to evaluate a candidate may come from another model’s policy. The candidate is therefore evaluated on a distribution of contexts that differs from the one its own behavior would produce. Each response changes what happens next: a model introduces a plot development, the user reacts, and that reaction becomes the model’s next input. After several turns, two models starting from the same context may be having very different conversations. Fixed-context comparisons miss this feedback loop and do not directly measure performance under the candidate’s own policy.
User-level A/B testing captures those consequences. Users experience the model over time, and we measure engagement, retention, and other product outcomes. But these experiments require traffic, coordination, and an observation window appropriate to the outcome. Testing every checkpoint this way would commit substantial user exposure before we had enough evidence to prioritize the candidates. We needed a practical intermediate step: give a model enough room to shape the interaction, measure the resulting behavior, and make that experiment easy to run repeatedly.
Elo comparisons
Which reply does the user prefer?
A telescope points at an empty sky. A notebook lies open to today’s date.
The dome creaks. Beside the telescope sits a cup of tea, still warm.
User-level A/B tests
How do users behave over time?
Evaluating the experience a model creates
In our product’s current form, a conversation is the smallest complete unit of content. Its quality depends on how the interaction develops across multiple exchanges. The motivation is similar to using watch time and video completion rate in recommendation systems: how long someone watches and whether they finish a video help capture their experience with that unit of content beyond the initial click. For conversations, we likewise want to capture the user’s experience across the interaction, including whether it sustains their interest and participation.
A model, its prompt, or the surrounding harness can change the direction of that experience, so the candidate configuration needs to remain in place across multiple exchanges for us to evaluate its effect. Conversations are open-ended and have no fixed endpoint equivalent to the end of a video. We therefore measure how the interaction develops and whether users continue participating, using bounded sessions as the unit of measurement.
We also needed to make this evaluation easy to repeat. Researchers should be able to compare checkpoints and presets, control their exposure to users, and inspect results without a new engineering handoff for every experiment. Session AB combines these requirements: a consistent model assignment over a sustained interaction, supported by a shared workflow for running and reviewing experiments.
A conversation can persist for days, so we still needed a bounded unit of measurement. We use sessions, with explicit rules for inactivity and other boundaries such as resets or model changes. For example, a 30-minute inactivity timeout separates sessions after a sufficiently long gap. It does not limit an active session to 30 minutes. The timeout determines how much of the user’s return behavior belongs to the same observation.
A single user can generate multiple sessions in a day. That gives session-level evaluation more observations than a user-level experiment over the same period, helping us collect feedback faster.
The system uses the conversation as its routing key and constructs bounded sessions from the resulting events. Stable assignment lets the candidate influence the trajectory of the interaction. The scorecard then compares behavior across the baseline and candidate strategies.
The main controls make those choices explicit:
- Candidates and baseline: which checkpoints, presets, or configurations to compare.
- Eligible traffic and capacity: which interactions can participate and how much exposure the experiment receives.
- Session boundaries: how inactivity and changes to the conversation define the observation window.
- Sample budget and stopping settings: how much evidence to collect and when to review the experiment.
Experiment setup
Strategies
Eligibility
Experiment parameters
We measure several aspects of the interaction. Generation turns describe how much activity occurred. Net depth subtracts regenerations, helping distinguish additional outputs from forward progress. Session span measures elapsed time between the first and last event. Regeneration behavior and latency provide additional context.
Those metrics need to be read together. More turns could partly reflect repeated attempts to obtain a satisfactory response. A longer session span includes gaps between actions and is not a direct measure of active reading time. For our storytelling products, sustained engagement is useful evidence, but the scorecard should help us understand how that engagement changed.
The workflow brings configuration, traffic assignment, event collection, and analysis into one place. A researcher can put a candidate into an experiment and inspect its results through the same interface.
From model deployment to experiment submission
We provide skills that handle model deployment and experiment submission, so researchers can run this workflow through their personal agents. All experiment submissions are made by those agents. This makes it easier to move from a candidate checkpoint to an online experiment and repeat the process as new candidates become available.
Running more experiments should not mean exposing more users to unchecked models. Candidates pass a safety gate before reaching real users, and we monitor the percentage of total traffic involved in experiments, as well as the size of each individual test.
Exposure is also bounded by the session. Unlike persistent user-level A/B assignments, Session AB treats a conversation reset as a stopping condition: users can reset to leave the current experimental session rather than remain assigned to the candidate.
The platform also provides latency controls to help isolate the effect of model behavior from differences in response speed. Controlling latency reduces a source of confounding when comparing engagement across candidates.
Experiment console
| Configuration | Status | Sample progress | Traffic control | Monitoring |
|---|---|---|---|---|
022ebb33B · kaon-v3V1 · kaonai/kaon-l-v3-s200 | Running | Sample progress1% | QPSTarget rate | |
69ce8d94B · kaon-v3V1 · kaonai/kaon-j-v14-20 | Running | Sample progress3% | slotsConcurrency | |
d7c90c5dB · kaon-v3V1 · kaon-v3-exp01 | Completed | Sample progress100% | QPSTarget rate | N/A |
12777401B · kaon-v3V1 · kaonai/kaon-l-v4-s525 | Running | Sample progress7% | QPSTarget rate | |
ec01c832B · kaon-v3V1 · kaon-c-r4-v0.1-420 | Completed | Sample progress5% | slotsConcurrency | N/A |
Reading a real experiment
Consider one experiment launched in August. We compared four model variants from a training run against Kaon v3. Absolute sample sizes and metric values are omitted; results are reported as relative lifts.
Primary metrics
Depth lift
Session span lift
Net depth
| Metric | Relative change |
|---|---|
| Generation turns per session | +6.8% |
| Net depth | +6.8% |
| Session span | +9.0% |
The reported lift intervals were +4.2% to +9.4% for generation turns and +6.2% to +11.7% for session span. Both metrics met the platform’s significance criteria. The rest of the scorecard adds useful context. The regeneration-rate estimate was slightly higher, with an interval spanning zero change. Median time to first token also increased.
This is the kind of result that supports a concrete next decision: the candidate shows stronger session engagement, with a latency tradeoff to investigate before broader exposure. It gives us evidence to prioritize further testing.
Evaluating over a hundred variants a week
In practice, Session AB can provide directional results within hours and estimates of lift in about half a day, depending on traffic and effect size. This lets researchers identify promising candidates quickly while continuing to collect evidence.
Session AB makes evaluation a routine part of model development. Across four complete weeks, we ran 388 experiments with samples and evaluated 521 distinct challenger model configurations, averaging 97 experiments and 136 challenger configurations per week. Configurations can include different checkpoints, prompts, or serving setups, and some are evaluated in more than one week. This scale lets researchers explore broadly and use session-level evidence to prioritize candidates for larger user-level A/B tests.
How well does Session AB track user-level A/B?
Throughput only matters if the feedback is useful. We therefore examined how Session AB results corresponded to outcomes in user-level A/B tests. Across 30 matched model arms from three A/B batches, the correlation between Session AB session-span lift and first-day active chat-time lift in user-level A/B was:
| User population | Pearson correlation |
|---|---|
| Existing users | 0.93 |
| New users | 0.71 |
Session AB vs. user-level A/B
Existing users
r = 0.93New users
r = 0.71The relationship remained strong when we removed each A/B batch in turn. For existing users, the correlation ranged from 0.92 to 0.95; for new users, from 0.69 to 0.73.
This retrospective analysis supports the role we wanted Session AB to play: screening candidates using a short-term engagement signal that corresponds closely to engagement in broader experiments. It does not establish the same relationship for retention, which still requires its own observation window and validation.
A direct comparison with Elo needs the same models and A/B outcomes on both sides. In our available Chat Elo export, matching configuration IDs against the broader A/B–Session AB set, including normalization of the kaonai/ prefix, recovered just one shared candidate identity. That is insufficient to estimate an Elo–A/B correlation. The evidence here therefore establishes a strong retrospective association between Session AB and short-term A/B engagement; it does not quantify how much more predictive Session AB is than Elo.
A faster path from ideas to evidence
The design principle behind Session AB is simple: the unit of evaluation should reflect the experience we want to improve. For conversational products, that means allowing a model to shape more than one response and observing how users interact with what it creates. Building infrastructure around that unit gave us a repeatable way to evaluate over a hundred model variants a week. It lets us explore more ideas, identify promising candidates, and bring better evidence into the larger A/B tests that guide product decisions.
