AI systems.
Engineered for
the real world.

Most AI pilots never reach production. I build the ones that do, and the platforms that keep them running.

Sree Teja Mutyala
Engineering Lead, AI & Distributed Systems

Data, intelligence, and infrastructure
designed as one system.
From the first architecture decision to production.Explore selected work

What teams
bring me in for.

For teams moving AI into a real product, improving what’s already running, or building the platform to support it.

Bring AI into production

Document intelligence, retrieval, and agent workflows built around your data, permissions, and users.

RAG · Agents · Human-in-the-loop

Make quality and cost visible

Evaluation and tracing that turn model, prompt, and retrieval choices into decisions you can defend.

Evaluations · Cost attribution · Observability

Build a platform you can operate

Cloud infrastructure, delivery pipelines, and observability designed for the people who run them.

AWS · Kubernetes · Distributed systems

Selected work

Engineering work at Iluminr, spanning the AI product and the infrastructure beneath it.

From documents to
usable intelligence.

A multi-tenant document platform needed to make complex files useful in chat, while respecting who could access them and what it cost to process them.

I built ingestion, structure-aware chunking, hybrid retrieval, and streaming agent workflows, with tenant isolation and permissions enforced in retrieval.

The engineering decisions

Instrument the pipeline before optimizing it. Per-stage tracing exposed the expensive extraction calls; reducing call volume, narrowing the entity schema, and using prompt caching cut inference spend. Human-in-the-loop workflows let the agent pause for user input.

Document intelligence · Retrieval · Agent workflows
Illustrative knowledge graphSynthetic example

The supplier connects a source document, a service, a region, and supporting evidence.

Select a node to explore its connections.

Making inference cheaper
and proving quality held.

Cheaper model calls are only useful if the report still does its job. I built an evaluation bench before changing the report-generation pipeline.

Repeatable comparisons, an LLM judge, cost attribution, and promoted baselines made it possible to evaluate model, retrieval, and prompt changes together.

What the result means

The measured 39.6% reduction combined three winning changes, with quality inside the evaluation’s baseline band. Failed experiments and corrected conclusions stayed in the record. This is a result from that evaluation, not a guarantee across other workloads.

LLM evaluation · Quality gates · Inference economics

Relative report-generation cost

Baseline100%
Combined changes60.4%

39.6% lower cost.
Quality within the baseline band.

Measured in the report-generation evaluation.
Relative cost, with baseline normalized to 100%.

The platform behind
the product.

As the product grew, the platform needed a repeatable way to deploy, scale, and operate across regions.

I led the move from ECS to EKS and built infrastructure as code, GitOps delivery, autoscaling, and a shared observability stack, then delivered production rollouts in the US, Australia, and Europe.

Operating beyond the rollout

Ownership continued through fleet upgrades and cross-region content distribution. Phased release playbooks, disruption budgets, idempotent operations, and stale-data fallbacks made failure handling part of the design. Traces, metrics, logs, and alerts supported day-to-day operation.

AWS / EKS · GitOps · Platform ownership

Three regions.
One operating discipline.

US · Australia · Europe
Infrastructure as code, delivery, and observability.

An engineer who
owns the whole picture.

I’m Sree, an Engineering Lead based in Hyderabad, India. My work connects AI product decisions with the systems, teams, and operating practices that make them hold up.

I’ve taken a product from concept to market, led 13 engineers, and built developer tools that make a team’s knowledge easier to use. I care about clear handoffs and helping other engineers grow. That work earned a Best Mentor Award and a CEO’s Award at Xyenta.

More about my experience

Experience

2024 – Now

Iluminr

Engineering Lead Jul 2026 – present

Senior Software Engineer Nov 2024 – Jun 2026

2024

Velocity Clinical Research

Full Stack Engineer

2021 – 2024

Xyenta Solutions

Lead / Founding Engineer

Previously Software Engineer

2018 – 2021

The foundations

Full Stack Blockchain Developer, Nvest
Freelance Software Developer, 9and9

What are you
working on?

Production AI, platform decisions, or a problem you’re stuck on. I’m always glad to talk about this kind of work.