Skip to content
← All writingKaon AI / Research

Offline Evaluation Is Dead: How Consumer AI Learns from Millions of Real User Decisions

Once consumer AI products reach scale, offline evaluation isn't just "not good enough"—it's irrelevant to production decisions. This conclusion comes from running systems at tens of millions of users, processing billions of conversations and behavioral signals every month. We've seen this pattern repeatedly in real user environments: models that top public creative writing benchmarks often underperform in actual usage.

Once consumer AI products reach scale, offline evaluation isn't just "not good enough"—it's irrelevant to production decisions. This conclusion comes from running systems at tens of millions of users, processing billions of conversations and behavioral signals every month.

We've seen this pattern repeatedly in real user environments: models that top public creative writing benchmarks often underperform in actual usage. Meanwhile, models with unremarkable offline scores consistently drive better retention, deeper conversations, and higher continuation rates. In some cases, multiple benchmark top-5 models get outperformed by a model ranked significantly lower offline.

This isn't anecdotal. It's structural.

Offline benchmarks solve for discrete, bounded problems—tasks with clear boundaries and correct answers. Consumer AI faces something fundamentally different: continuous content consumption experiences that span multiple conversation turns and contexts. In entertainment and creative writing, there's no objective ground truth, only preference. Benchmarks designed for discrete tasks can produce scores, but those scores don't reliably map to how people actually behave at scale.

Offline evaluation is failing—not because benchmarks are poorly designed, but because they're fundamentally misaligned with consumer AI.

As model costs continue dropping in 2026 and user expectations mature rapidly, the real product edge won't be who has the best demo—it will be who can iterate faster, and who can make the right call on which model deserves to stay. In this context, evaluation is no longer a supporting system—it's the core bottleneck determining product ceiling. Yet in consumer AI, there's still no mature, scalable evaluation methodology.

The problem isn't whether model responses are "accurate." The problem is defining what success even means.

The Fundamental Mismatch

At its core, an offline benchmark is a set of artificially constructed tasks: tasks defined by humans, graders designed by humans (or LLMs), success criteria locked in before experiments begin. This might work in constrained scenarios. But in large-scale consumer AI, pre-defining success is inherently unsustainable.

Content products learned this lesson years ago. As Eugene Wei points out in his analysis of TikTok, the system isn't trying to decide whether content is "good." It continuously observes behavior, infers each user's preference distribution, and ranks accordingly. In that setup, behavior is the only signal that scales, and any evaluation framework that tries to define "good content" upfront collapses in front of real users.

In consumer apps, user behavior is the only ground truth.

This is exactly why offline eval systematically fails in consumer contexts: it optimizes for "answers that look correct," while users reward "experiences that feel right." When evaluation systems decouple from real user behavior, they can produce scores, but those scores can't support any critical decisions.

We saw a direct example on Emochi (our creative fiction product): among models ranked top-5 on a public Creative Writing benchmark, several showed lower user retention than models ranked much lower offline. That's not a one-off. It's a warning: when "scoring criteria" and "user preferences" aren't the same thing, offline scores become an illusion.

Offline Benchmark Score vs. Online Retention

Offline Benchmark Score vs. Online Retention

For consumer AI, any evaluation system that can't be driven by real user online behavior isn't "incomplete"—it's irrelevant. Scalable online evaluation infrastructure based on real user feedback is no longer optional. It's the only path forward.

Evaluation Only Scales When It's Closed-Loop

In consumer AI, model evaluation is no longer a standalone step—it's a continuously running closed-loop feedback system. Models get deployed, used, compared, filtered, trained, then return to real user environments for validation. The core of this loop: all critical evaluation signals come from real user online behavior.

In production, the loop converges to something like this:

models → relative comparison (Elo) → decision → reward modeling (RM) → signal → models

Closed-Loop Evaluation System

Closed-Loop Evaluation System

Everything runs on real user data. Relative comparison continuously harvests preference signals. Those signals become operational decisions—what to ramp, keep, or cut. And the same feedback is modeled into product-specific rewards that drive the next training cycle.

Within this system, Elo and RM aren't parallel evaluation tools—they're system components operating at different layers, solving different problems:

Elo solves for scale. RM solves for how evaluation enters training.

When these components unify within the same closed loop, evaluation stops being a step that "validates whether a model is good." It becomes infrastructure driving continuous model evolution. Any evaluation system not built on real user behavior will ultimately optimize for a product that no longer exists.

Ranking Models by User Preference, Not Static Scores

At this stage, the primary constraint on evaluation systems isn't statistical precision—it's throughput and time-to-decision.

Traditional A/B testing assumes three premises: stable evaluation targets, sufficiently long experiment windows, and slowly changing model sets. In real consumer AI production environments, nearly all these premises fail. New models might join on a daily or even hourly basis. Before a traditional A/B experiment converges, candidate models are often already obsolete.

Relative ranking systems (like Elo, or more precisely, TrueSkill) happen to fit these constraints. They're designed to tolerate noise, incomplete matchups, and continuously changing opponents—characteristics highly aligned with online traffic scenarios.

We don't use Elo to find a "globally optimal model." We use it to solve a more engineering-focused problem: given current traffic and time windows, which models can be confidently eliminated.

More importantly, we don't treat ranking results as "absolute scores"—we treat them as online estimates with uncertainty. TrueSkill uses (μ, σ) to explicitly express "current strength" and "confidence level." This allows the system to pursue stable relative ordering even when noise is unavoidable and preferences continuously drift—rather than chasing a seemingly precise but production-unreliable single-point score.

To keep ranking controllable over long horizons, we run it like infrastructure, not a one-off experiment:

  • Basic sample-quality filtering (e.g., removing unstable matchups with too few turns)
  • Robustness checks via merging and shuffling to reduce sensitivity to matchup order
  • A stable set of reference models as anchors, to stabilize scale, cold-start new candidates, and reduce drift as the pool churns

Elo/TrueSkill is always a filter, never a final truth. Its job is to cheaply rule out bad bets fast, and reserve heavier evaluation bandwidth for the candidates that survive. In our current traffic, we can typically eliminate obviously mismatched models within hours, using hundreds of thousands of comparison samples—then focus deeper A/B and RM resources on the top pool.

Elo/TrueSkill Convergence: Early Volatility, Then Stabilization

Elo/TrueSkill Convergence: Early Volatility, Then Stabilization

Turning User Behavior into a Learning Signal

If evaluation can only tell us "which model is better" but can't influence model evolution itself, it ultimately remains just a post-hoc explanation system, not part of the production system.

In consumer AI, iteration speed sets the ceiling for what evaluation is worth. The loop only becomes real when evaluation signals flow back into training continuously and reliably—so model selection, optimization, and real user preference actually close.

This feedback pipeline accumulates hundreds of millions of implicit labels and large volumes of pairwise preference samples daily, giving reward signals the density and coverage needed for sustainable training.

This is exactly RM's (Reward Model) role in the entire system.

Unlike generic reward models, Emochi's RM doesn't start from abstract preferences or human annotations—it's built directly on real user behavior. The system doesn't assume users will explicitly tell us "what a good model is." It assumes users will continuously reveal true preferences through behavior.

These behavioral signals include but aren't limited to: whether users continue conversations, next-turn query length, dwell time on model responses, explicit like/dislike, and actual choices during model matchups. These signals may be noisy in single interactions, but when aggregated at scale, they're repeatedly validated as highly correlated with long-term retention and engagement.

The key is that these signals aren't used in isolation. RM's function is to unify user feedback from different evaluation system layers into a learnable reward space, enabling model training to directly align with real product goals rather than some static offline metric.

RM isn't a one-time trained model—it's a continuously calibrated system. As user behavior distributions, product forms, and model capabilities change, reward signals themselves will drift. The system design goal isn't to eliminate this drift—it's to maintain long-term correlation between rewards and real user value under inevitable drift.

When evaluation signals, model decisions, and training feedback unify in the same closed loop, model optimization no longer depends on offline assumptions—it's directly constrained by real user behavior. This enables the entire system to continuously self-correct at scale rather than relying on periodic manual calibration.

In this sense, RM isn't an auxiliary module of the evaluation system—it's the final piece that makes consumer-scale online evaluation a "platform": it moves evaluation beyond the selection layer into the driving layer of model evolution.

This closed loop didn't start from theoretical derivation. It gradually took shape in Emochi's real user environment, accompanying growth in model scale, traffic, and product complexity. Through long-term operation, we repeatedly observed: when evaluation decisions are truly handed over to user behavior, the correlation between offline scores and core business metrics naturally recedes to secondary importance.

Trade-offs and Failure Modes

This scalable online feedback loop wasn't designed to solve all evaluation problems. Its goal is very specific, which means it makes clear trade-offs.

First, this infrastructure isn't suitable for all product stages. When product scale is still small (e.g., DAU < 50k), user behavior signals are too sparse and noisy to support stable online ranking and decisions. In such cases, human evaluation or offline analysis may still be more efficient paths. But once products cross this scale threshold, continuing to rely on offline evaluation is no longer the "safe choice"—it's outdated system design. For AI systems running in real user environments, if evaluation signals don't come from real user behavior, they will inevitably systematically diverge from product goals.

Second, this isn't an evaluation system for answering "how good is a model in absolute terms." The system doesn't try to produce cross-time, cross-scenario, reproducible absolute scores, nor does it pursue one-to-one alignment with traditional benchmarks. The only question it cares about: given current users, current distribution, current product form, which models deserve continued scaling and evolution.

Third, this system doesn't assume user feedback is clean, stable, or interpretable. On the contrary, it accepts real-world constraints from the start: user behavior is noisy, preferences drift, and product forms themselves continuously change. The system design goal isn't to eliminate these uncertainties—it's to continuously reduce decision risk when uncertainty is unavoidable.

Finally, this isn't a "one-time evaluation" solution. It can't provide definitive conclusions before models enter real usage scenarios, nor can it replace offline analysis during research stages. Its premise is that models have already been deployed, already used by real users, and evaluation itself needs to become part of product and training workflows.

Precisely because these boundaries exist, this system can operate long-term at real consumer AI scale. It abandons pursuit of "static truth" in exchange for reliable support for "continuous decisions." It abandons illusions of perfect evaluation in exchange for a production-grade platform capable of continuous self-correction.

This is what we've built:

Not an evaluation tool, but feedback infrastructure that can evolve alongside models and products.

Interested in this work?

We’d love to hear from you. For research questions or collaboration, reach out to alex@kaonlabs.com. If you’d like to help build what’s next, explore our open roles.

Continue reading

Online Evaluation at Scale: How We Evaluate 100+ Model Variants per Week

Reading a roleplay model before it speaks: what differs after post-training

Online Autoresearch: Running Autoresearch with Real Online Feedback

Build the medium while
it is still becoming.

Join the team building consumer AI across product, research, and infrastructure.

View open roles