Skip to content
All posts
ResearchJuly 2, 2026

Evaluating what matters: benchmarks for real work

M

Meera Krishnan

Research Lead

Model leaderboards make great headlines and poor purchasing decisions. A model that tops an academic benchmark can still fail at your specific work — and a cheaper model may quietly excel at it. The only benchmark that matters is the job in front of you.

Task-grounded evaluation

Our research group builds evaluations from real workflows: does the tutoring system actually improve a learner's next attempt? Does the documentation assistant produce notes clinicians sign without heavy edits? Does the support assistant resolve tickets or merely respond to them? Each benchmark is a job description, not a trivia quiz.

What we're finding

Early results keep teaching us humility. Bigger models don't uniformly win — on structured extraction tasks, well-prompted smaller models with good retrieval routinely match models ten times their price. Multi-agent decomposition helps on long-horizon tasks and actively hurts on short ones. Indian-language performance varies far more across models than English benchmarks suggest, which matters enormously for our market.

We'll publish these benchmarks openly as they mature. Measurement is a public good — and honest measurement is how the whole field gets more practical.