Zepto Adopts Evaluation-First Multi-Agent Systems to Scale Customer Support
Indian quick-commerce platform Zepto has partnered with Databricks to implement an evaluation-first framework for its multi-agent customer support system.

Key takeaways · 3
- 01
Zepto's AI system handles over 100,000 support tickets daily.
- 02
A 1% error rate at scale causes thousands of daily failures and revenue leakage.
- 03
Multi-step agent workflows require evaluation at every stage to prevent hidden failures.
Scaling AI Support
Zepto, an Indian quick-commerce platform operating in over 60 cities, uses a multi-agent artificial intelligence system to manage customer support. [1] The platform processes more than 100,000 support tickets daily. [1] Customer behavior shifts, expansion into new categories like electronics and beauty, and seasonal events like Diwali drive spikes in ticket volume. [1] At this scale, a one percent error rate results in thousands of bad outcomes and actual revenue leakage each day. [1]
Addressing the Assurance Gap
To maintain reliability, Zepto partnered with Databricks to adopt an evaluation-first approach to building, testing, and operating its agents. [1] The company identified an "assurance gap" in its multi-step workflows, noting that failures can occur during intent classification, knowledge retrieval, or tool calling rather than just in the final generated response. [1] Before implementing this evaluation framework, failures remained invisible until customers complained, internal errors were hidden by final answers, and fixes were slow. [1]
What it means
The transition from simply shipping agents to adopting an evaluation-first methodology highlights a critical maturation point for enterprise AI operations. When multi-agent systems reach the scale of hundreds of thousands of daily interactions, traditional end-point monitoring fails to capture intermediate logic errors in reasoning or retrieval. Zepto’s reliance on Databricks and MLflow indicates that infrastructure tooling is shifting from basic deployment capabilities toward robust, step-by-step observability. As platforms expand across languages and product categories, managing the assurance gap becomes a primary operational requirement rather than an edge case. What the sources don't address: How the evaluation-first framework specifically impacts the latency of customer support responses during peak volume events.
As companies scale AI agents in production, traditional monitoring falls short. Organizations must build evaluation frameworks that inspect intermediate steps like intent classification and tool execution.
Why it matters
Put this to work — one session a day, built for your industry.
Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.
Start freeHow this developed
9 September 2026
Zepto Adopts Evaluation-First Multi-Agent Systems to Scale Customer Support
9 September 2026
Event created from source cluster.