Bring AI into production
Document intelligence, retrieval, and agent workflows built around your data, permissions, and users.
RAG · Agents · Human-in-the-loopMost AI pilots never reach production. I build the ones that do, and the platforms that keep them running.
Sree Teja Mutyala
Engineering Lead, AI & Distributed Systems
For teams moving AI into a real product, improving what’s already running, or building the platform to support it.
Document intelligence, retrieval, and agent workflows built around your data, permissions, and users.
RAG · Agents · Human-in-the-loopEvaluation and tracing that turn model, prompt, and retrieval choices into decisions you can defend.
Evaluations · Cost attribution · ObservabilityCloud infrastructure, delivery pipelines, and observability designed for the people who run them.
AWS · Kubernetes · Distributed systemsEngineering work at Iluminr, spanning the AI product and the infrastructure beneath it.
A multi-tenant document platform needed to make complex files useful in chat, while respecting who could access them and what it cost to process them.
I built ingestion, structure-aware chunking, hybrid retrieval, and streaming agent workflows, with tenant isolation and permissions enforced in retrieval.
Instrument the pipeline before optimizing it. Per-stage tracing exposed the expensive extraction calls; reducing call volume, narrowing the entity schema, and using prompt caching cut inference spend. Human-in-the-loop workflows let the agent pause for user input.
The supplier connects a source document, a service, a region, and supporting evidence.
Select a node to explore its connections.
Cheaper model calls are only useful if the report still does its job. I built an evaluation bench before changing the report-generation pipeline.
Repeatable comparisons, an LLM judge, cost attribution, and promoted baselines made it possible to evaluate model, retrieval, and prompt changes together.
The measured 39.6% reduction combined three winning changes, with quality inside the evaluation’s baseline band. Failed experiments and corrected conclusions stayed in the record. This is a result from that evaluation, not a guarantee across other workloads.
Relative report-generation cost
39.6% lower cost.
Quality within the baseline band.
Measured in the report-generation evaluation.
Relative cost, with baseline normalized to 100%.
As the product grew, the platform needed a repeatable way to deploy, scale, and operate across regions.
I led the move from ECS to EKS and built infrastructure as code, GitOps delivery, autoscaling, and a shared observability stack, then delivered production rollouts in the US, Australia, and Europe.
Ownership continued through fleet upgrades and cross-region content distribution. Phased release playbooks, disruption budgets, idempotent operations, and stale-data fallbacks made failure handling part of the design. Traces, metrics, logs, and alerts supported day-to-day operation.
US · Australia · Europe
Infrastructure as code, delivery, and observability.
I’m Sree, an Engineering Lead based in Hyderabad, India. My work connects AI product decisions with the systems, teams, and operating practices that make them hold up.
I’ve taken a product from concept to market, led 13 engineers, and built developer tools that make a team’s knowledge easier to use. I care about clear handoffs and helping other engineers grow. That work earned a Best Mentor Award and a CEO’s Award at Xyenta.
More about my experienceEngineering Lead Jul 2026 – present
Senior Software Engineer Nov 2024 – Jun 2026
Full Stack Engineer
Lead / Founding Engineer
Previously Software Engineer
Full Stack Blockchain Developer, Nvest
Freelance Software Developer, 9and9
Production AI, platform decisions, or a problem you’re stuck on. I’m always glad to talk about this kind of work.