Raindrop Hits $50M Total To Catch AI Agents Failing Before Users Notice

Raindrop Zubin Singh Koticha (CEO), Ben Hylak, and Alexis Gauba
Image credit: Raindrop linkedin
Raindrop, a San Francisco based AI agent reliability company, has raised a Series A round led by CRV that brings its total funding to 50 million dollars, alongside the launch of a new product designed to catch agent failures before they ever reach a live user.
Existing investors Lightspeed Venture Partners and Y Combinator returned for the round, joined by angel investment from senior researchers and leaders at OpenAI, Anthropic and Thinking Machines. The raise follows a 15 million dollar seed round led by Lightspeed in December 2025, meaning Raindrop has added roughly 35 million dollars in funding in under a year, a pace that tracks closely with how quickly its customer base has grown over the same period.
Raindrop was founded in 2023 by chief executive Zubin Koticha, Alexis Gauba and chief technology officer Ben Hylak. Koticha and Gauba previously co‑founded Opyn, a decentralised finance options platform that processed more than 15 billion dollars in trading volume before being acquired by Coinbase, while Hylak spent four years on Apple's Human Interface team working on visionOS. According to the founders, Raindrop began as an internal tool they built to debug their own coding agent, only becoming a company once they realised the debugging problem they had solved for themselves was one every team building agents was going to run into.
That problem has grown more urgent as AI agents have taken on longer, more autonomous and more consequential work. Koticha has described the core failure mode Raindrop is built to catch in fairly stark terms, noting that agents now run for hours, call thousands of tools, and handle real money, real health data and real customers, and that when an agent fails, it tends to do the wrong thing confidently and at scale until a person happens to notice. Traditional monitoring tools built for conventional software were never designed to catch that kind of failure, since they typically expose only surface level metrics such as latency, token usage or generic toxicity scores, leaving engineering teams blind to the deeper behavioural problems that actually cause agents to go wrong in production.
Raindrop's core product addresses that gap by detecting semantic anomalies directly inside production agent traffic, using small, custom models adapted to the specific shape of each customer's AI product rather than relying on generic, one‑size‑fits‑all signals. That approach surfaces default signals such as user frustration alongside product‑specific ones customers define themselves, including patterns like an agent becoming stuck in a repeating loop or complaints tied to a specific part of a product's interface, tracking incident rates across millions of events and issuing Sentry‑style alerts when something shifts. When behaviour changes, engineering teams can see what changed, when it started and which users were affected, backed by real examples pulled directly from production traffic rather than a single flagged data point with no surrounding context.
Reid Christian, general partner at CRV, has framed the firm's decision to lead the round around a structural difference between agents and the software categories his firm has historically backed, arguing that agents are fundamentally different from traditional software because they are highly capable, autonomous and non‑deterministic, and that Raindrop's approach of treating agent failure as a detection problem, the way a security company would treat a network intrusion, is the right frame for a category of software that cannot be fully specified or tested in advance the way conventional applications can. Bucky Moore of Lightspeed, who backed Raindrop's seed round, has pointed to the company's customer trajectory as validation of that thesis, noting that barely a year after the seed round, Fortune 100 companies are now running Raindrop directly on their agent traffic, a jump from the startup‑heavy customer base, including Vercel, Framer, Clay and Speak, that the company built its early track record on.
Alongside the funding announcement, Raindrop launched Simulations, currently in research preview, extending the company's detection approach earlier into the development process rather than limiting it to catching problems after an agent is already live. Simulations automatically recreates realistic, fully mocked versions of every service and database an agent depends on, letting engineering teams replay thousands of historical production traces against a new version of their agent on every single pull request, before any change ships. Hylak has said that testing agent changes this way is among the most effective methods the team has found for catching regressions ahead of deployment, while acknowledging that building simulations robust enough to survive changes to an agent's own underlying harness, such as adding a new tool, is a genuinely difficult technical problem in its own right.
Raindrop has drawn a direct parallel between Simulations and testing methodologies already used inside frontier AI labs, pointing to OpenAI's published research on deployment simulation, which regenerates responses to de‑identified production conversations with a candidate model to estimate misbehaviour rates before a release goes out, and to Anthropic's use of synthetic universes to train and stress‑test its own agents internally. Raindrop's pitch is that Simulations brings a version of that same frontier‑lab testing rigour to any team building agents, rather than leaving that level of pre‑deployment testing available only to the small number of labs with the internal resources to build it themselves.
The new funding will go toward accelerating Raindrop's underlying anomaly detection research, scaling the core product to a broader set of enterprise customers, and continuing development of Simulations, which the company expects to move from research preview to general availability over the coming month. Raindrop's rise reflects a wider shift taking place across AI infrastructure investment, as capital moves from purely capability focused AI tooling toward the reliability and safety layer needed to deploy that capability responsibly in high‑stakes settings such as healthcare, finance and logistics, where an agent's confident mistake can carry consequences well beyond a frustrated user. Whether Raindrop's detection‑first approach continues to scale as agents grow more autonomous and harder to fully specify in advance, rather than being overtaken by reliability tooling built directly into the frontier model providers' own platforms, is likely to shape how durable a position the company can hold in what is already becoming a more crowded corner of AI infrastructure.
Topics
Sources
Stay informed
Startup news in your inbox
Get important funding rounds, founder stories, and startup updates.
No spam - only important startup updates.





