devopsdays Portland 2026

Ankit Agarwal

Ankit Agarwal is a Senior Software Engineer at TikTok, Inc. in Seattle, Washington, with over ten years of specialized experience across four of the most technically demanding organizations in the global technology industry. Ankit began their career at SAP SE, contributing to backend itinerary services infrastructure within the enterprise travel division of a company serving over 400,000 customers across 180 countries. Ankit then joined Amazon, where they engineered real-time fraud detection systems on the buyer fraud prevention team - work that directly protected the financial integrity of one of the world's highest-volume transactional platforms, serving hundreds of millions of customers. At Microsoft, Ankit served as a software engineer on the Azure Container Registry team, contributing to the cloud-native artifact infrastructure that underpins deployment pipelines for enterprises globally. Ankit currently holds a critical role on TikTok's Decisions Platform team, responsible for the real-time decision infrastructure serving over one billion users. Key contributions include leading the live migration of a production decision engine with zero behavioral change, designing a shadow and equivalence testing framework to validate correctness under live traffic, building configuration-drift detection systems, and pioneering the integration of AI tooling into production DevOps workflows. This decade-long trajectory across SAP, Amazon, Microsoft, and TikTok reflects a sustained record of original, high-impact engineering contribution at a level characteristic of a professional at the top of their field.


Session

09-09
11:20
30min
Swapping the Engine Mid-Flight: Migrating a Live Decision System with Zero Behavior Change
Ankit Agarwal

We replaced the engine behind a high-traffic, real-time decision system - while it kept running, with no user-visible change. This is the honest story of de-risking a high-stakes migration from a legacy rule engine to a new platform, and the techniques that made it survivable.
I'll walk through what actually worked: shadow/replay testing that ran old and new side-by-side on real traffic and diffed every response; a switchable flow that let us flip and instantly un-flip the traffic without a code change or deployment; automated configuration-drift detection between the two systems; and a staged rollout where rollback was a designed feature, not a panic button. You'll hear the decisions and tradeoffs behind each choice, the bugs that equivalence testing caught but code review missed, and the things we still got wrong.
If you're staring down a scary migration of something that can't go down, you'll leave with a concrete, reusable playbook and a more honest sense of what "zero downtime" really costs.

Main Track
Ballroom