Maarika Teose
Maarika Teose is a technology leader with over a decade of experience driving innovation across diverse industries. She is currently the Head of DevSecOps for a global automotive OEM, where she focuses on delivering secure, scalable software development platforms for next-generation Advanced Driver Assistance Systems. Previously, she led DevOps for Powell’s Books, and was a Release Manager at Jama Software. She is passionate about leadership and developing high performing teams. She is a proud mom of two kiddos and loves to dance.
Session
While companies increasingly adopt a hybrid multi-cloud approach to hosting and serving LLMs, there seems to be little in the way of a cohesive overview of what a "CICD pipeline" for LLMs looks like. I struggled for months to piece together an understanding of the basics, and I want to share what I've learned about what it takes to build and run LLMs in production.
I will cover hardware fundamentals, model optimization techniques like quantization, and the mechanics of serving high-performance AI inference, such as KV caching, batching, and disaggregation.
This is a field guide, not a how-to, and the implementation details will be different for each deployment. But if, like me, you’ve struggled to understand how to “DevOps” LLMs, I hope this talk will help to orient you so you can build your own mental model.
