The Long Way Home: How Evals 2.0 Came Back to Glific

Sep 2026

Writing this while sitting at the airport, travelling from Dehradun to Kochi. It has already been a long journey. The flight has been delayed multiple times—first from 15:25 to 17:00 and now to 22:10 from Delhi. We’ll probably land sometime after midnight.

It has also been a long quarter. Week after week, we’ve been working on Evaluations 2.0 in Glific, and now that it is finally out, we’re heading into a dev sprint in Kochi, followed by a team sprint in Bangalore.

So, a long quarter is ending with a long stretch of team events. And right after that, I’m heading for a trek to reset before the next quarter begins.

To kill some time at the airport, I was rewatching Nolan’s The Odyssey. At its heart, The Odyssey is one of the oldest stories about leaving home and finding your way back.

That made me think it might be interesting to structure this blog around some of those ideas—departure, the journey, the challenges along the way, and eventually, the return.

Ithaca, the home

I started as a developer writing Elixir code for Glific in January 2020 (yeah, I know—a long, long time ago). Over the years, I got to see both the product and the team mature into what they are today.

In January 2025, I moved from Glific to Kaapi, the API-first AI platform we were building at Tech4Dev for NGOs. I didn’t really leave the org. I just moved to a different product, a different codebase, a different language, and a different set of problems.

Meanwhile, Glific kept moving forward, with new team members taking it to new heights. Every quarter and during team sprints, I kept hearing about the new features being added and the new organizations joining the platform. Today, more than 250 NGOs are using Glific.

At Kaapi, we stayed focused on AI, starting with use cases from Glific. At the beginning of this year, we added an evaluations feature to Kaapi and started working closely with around 10 NGOs to help them run evaluations for their Glific bots.

More on evaluations in Kaapi in this blog:
https://projecttech4dev.org/ai-evaluation-from-seems-good-to-scores-good/

The next step was to use Kaapi’s APIs to bring evaluations directly into Glific.

Setting sail: Evals v1.0

In May 2025 we shipped the first version of evals in Glific. Upload a golden set of questions and answers, run the bot against them, get a score. Under the hood it used cosine similarity, and for the richer LLM-as-judge checks, it leaned on Langfuse.

It looked like a clean departure. Launch spike in June.

MetricValue
Eval runs by partner NGOs129
Orgs that ran evals30

Thirty orgs tried it. Very few came back. Two things sank us, and both were monsters we didn’t spot on the chart.

Calypso’s lotus: a score you forget to act on

In Nolan’s *Odyssey*, there’s a lotus-eating scene on Calypso’s island. On the shore, Odysseus eats the flower and slowly forgets Troy, then his crew, then his wife and son. Seven years pass. He isn’t dead. He just stops moving.

That was our cosine score.

What we saw was that most NGOs would run an eval and get 0.43. Then 0.78. Then 0.72.

But what do you do with 0.72? Nothing.

You look at it, nod, and go back to editing the prompt based on gut feel. The number doesn’t tell you what to fix or where to look. Teams would run an eval once, feel vaguely informed, and never come back to it.

The lotus.

For most teams, evals became something they ran before taking a bot live in the field. What we wanted instead was for evals to become a continuous habit: whenever a configuration change could affect the assistant’s behaviour, run the evals again and make sure everything still works as expected.

The Underworld: listening to the dead runs

Before he can get home, Odysseus has to go to the underworld and listen. Odysseus at edge of world, three ghosts: Sinon (liar of Troy), Agamemnon (won war, murdered at home within a week), Tiresias (which side of strait, what not to touch, only you return). “You don’t go to the dead for encouragement. You go for the map.”

Sinon = past usage data
Cosine scores from previous runs, usage patterns, how many NGOs used the feature, how often they used it, and when they stopped.

Agamemnon = internal experience
Feedback and ideas from our team, especially people who had worked closely with NGOs. Their experience helped shape Evals v2.

Tiresias = NGO feedback
NGOs told us what was missing in v1.0 and what they needed to make an informed decision.

Together, these gave us the spec.

Not “better metrics.” Three specific verdicts, each pointing to a specific fix, shown on the same screen where you edit the prompt. And fast enough that you would run it again before losing interest.

Scylla and Charybdis: two failures, one number

Set in Strait of Messina. Scylla built practical: massive body anchored in cliff, six serpentine necks ending in articulated jaw-heads, tiny eyes, rock-and-scale skin blending into cliff. Crew hugs cliff side to dodge Charybdis whirlpool. Scylla attacks in territorial rage, yanking six men off deck. Stunt team used ratchet rig to physically pull stuntmen off boat. Odysseus loses six men, ship survives.

Why it fits “two failures, one number”

Cosine score gave one number. Behind it sat two very different failures with no way to tell apart:

  • Scylla side: answer partially wrong. Facts missing from knowledge base, or model hallucinated past documents. Picks off individual answers.
  • Charybdis side: prompt itself broken. Wrong language, wrong tone, ignored scope. Swallows whole run.

And the strait was slow. A run took around 45 minutes. Long enough to lose the thread of what you were testing.

NGO sees 0.72. Can’t tell which monster bit. Fix KB? Fix prompt? Fix golden set? Guessing. Like steering strait with a compass that only reads “somewhat bad.”

The Cyclops: Glific codebase

Before we really got started, there was a fair bit of discussion around the five-year-old frontend codebase. For both me and Ayush, this was largely unfamiliar territory, and the codebase felt a bit like a giant Cyclops: intimidating, opinionated, and not something you casually walk into without understanding what you are dealing with. A lot of the early conversation was around how much we should change versus how much we should preserve, especially around MUI, theming, component structure, and whether we should continue extending the existing system or move toward something more flexible.

There were also long discussions in the Glific server around standardization, visual identity, migration effort, and avoiding another round of component bloat. Eventually, instead of trying to solve the entire frontend architecture at once, we narrowed things down into something much more practical: a clear list of components we would create, components we would reuse

The Sirens: things we had to hear and not steer into

Every feature has songs that sound like add-ons/must have. A few we tied ourselves to the mast for:

We started building evals 2.0 without any long UI/UX research phase. We gave Claude the spec, got a working prototype, and let the team pull it apart. Build fast, fail fast, iterate. This produced around a hundred small feedbacks, and most were fixed; some were passed on, and parked a few.

  • Put this behind feature flag
  • A summary per eval run. The shape of 500 answers in one read.
  • Prompt iteration from the run. Eval results plus current config in, suggested prompt out. Run evals on new version so see how new AI suggested prompt works.
  • Cost and latency per run. Wanted, not in the UI yet.
  • Guardrails for gender assumptions and safety. Designed, next up.

We also fixed few issues that only surface when people started using evaluations and common code/functions that were quietly failing.

Building the raft: what Evals 2.0 actually is

Kaapi is the backend. Glific is the frontend. We built them in lockstep, one PR on each side for every feature, because the whole point was that the loop should feel native.

The judge: A native LLM-as-judge replaces cosine similarity. Every answer gets three checks on a 0 to 5 scale:

  • Ground truth, 50%. Is the answer actually correct against the golden Q&A?
  • Knowledge base, 30%. Is it supported by the uploaded documents, or invented?
  • Prompt adherence, 20%. Did it obey tone, language, scope, refusals?

Three verdicts, not one number: Low KB score means fix your documents. Low prompt score means fix your instructions. Low ground truth means your golden set or your model is off. Each check points at a different door.

The summary: One ring score, a verdict label (Good, Could improve, Needs improvement), the judge’s written rationale, and weighted bars showing how the number was built. You read it in a glance instead of scrolling a CSV.

The assistant workflow: Assistants now have proper draft, publish and live versions with semantic labels like 1.0 and 1.1. Every saved version carries its own eval score, so you can see whether a change helped. A Try It Out sandbox lets you chat with any version before it goes live. Models sync from Kaapi’s catalog automatically, tagged recommended, all, or deprecated.

Prompt suggestions: After a run, one click asks the judge to rewrite the weakest part of your prompt. Apply it, or restore the old one. You don’t need to be a prompt engineer to close the loop.

No more polling: Version builds, KB indexing, eval completion and prompt-improve jobs all push to the screen over subscriptions. The results appear when they’re ready. You don’t refresh.

Speed: 500 items in under 4 minutes. Down from 45.

The bow

Here is the part of the Odyssey I kept thinking about.

Odysseus comes home disguised as a beggar. Nobody recognizes him. The suitors have been in his hall for years. Penelope sets a contest: string the great bow and shoot through twelve axe heads. The young men can’t even bend it. The beggar strings it in one motion and puts the arrow through.

He proves who he is not by announcing it, but by doing the one thing only he could do.

For me and Ayush it was a different challenge as I hadn’t written Elixir in a long while and he was writing frontend code in 5 year repo for first time and that too this big frontend-heavy feature. Coming to Glific for this feature, we weren’t sure the bow would still bend. The codebase had moved on. Patterns had changed. There were new people who knew it better than I did.

Then the first PR went in. Then the few changes from frontend were merged, all this still behind a feature flag so the initial hesitation of it breaking things in production was not there. The bow still bent. Not because I was special, but because the years in Kaapi had taught me exactly what the backend needed to send and exactly what the frontend needed to hear. Leaving and returning gave me both halves of the stack in one head.

The feature launch was the arrow through the axe heads. Evaluation is now native in Glific. No feature flag to wait for. No separate tool. Your bot can test itself and suggest its own fixes on the same screen where you built it.

Onto the sunset

The story doesn’t end at the bow. It ends when Odysseus set onto new sail, new journey

For us, after the launch we started working with five NGOs, walking through their real bots, their golden sets, and their evaluation runs. That’s where the judge prompt gets tuned. That’s where guardrails get their first shape eventually. That’s where we find out whether “three verdicts instead of one number” actually changes what a program manager does.

Since Evals 2.0 launched in early September, we’ve started seeing early signs of adoption. In the first two weeks alone, there were 97 eval runs out of 240 total runs so far. We’ve also been hearing positive feedback from every NGO we’ve spoken with.

If you talk to NGO partners and hear “the bot gives weird answers sometimes” or “we set it up months ago and haven’t touched it,” that’s the cue. Tell them their bot can now test itself inside Glific.

The voyage took a while. We’re home now. But this is where the real work begins: helping NGOs become more confident in the assistants they’ve set up, and giving them a way to keep improving those assistants over time.

You may also like

Notes from Protsahan’s Girl Empowerment Centre

Driving Beneficiary Impact from Data: Lessons from 1000 Days Fund

Understanding the Fundraising Challenge for Grassroots CSOs