AKSHAY NAGAR

Senior Software Engineer • AI Engineer • Inference Systems

7+ years building production systems at scale (JPMorgan Chase) | Now focused on AI engineering, LLM systems & optimizing inference at scale

🤖
⚙️
Akshay Nagar
✨ Open to opportunities
📍 Utah, USA

ABOUT_ME

💼 Experience

JPMorgan Chase

Senior Software Engineer

~7 years building at scale

Distributed systems & platform

🤖 Focus

AI Engineering

LLM systems & RAG pipelines

Inference optimization at scale

Applied ML in production

🚀 Currently

Open to roles

Senior SWE / AI Engineer

Building startup projects

Fast, high-quality shipping

HOW_I_WORK

I'm an engineer who's shipped mission-critical software at one of the world's largest financial institutions—and I like owning problems end-to-end. Lately I'm channeling that same rigor into AI: building LLM-powered systems that are not just impressive in a demo, but fast, efficient, and reliable in production—squeezing latency and cost out of inference at scale.

🏦

Enterprise Scale at JPMorgan

JPMorgan Chase • ~7 Years

Designed and shipped production systems serving enterprise-scale traffic in a highly regulated environment. Owned services end-to-end with strong reliability, security, and compliance requirements.

Distributed Systems Production High Stakes
🧠

LLM & RAG Systems

AI Engineering • Hands-on

Building retrieval-augmented and agentic LLM applications—owning the full pipeline from retrieval and prompting to evaluation and guardrails. Focused on making AI output testable and observable.

RAG LLM APIs Evals
⚙️

Optimizing Inference at Scale

Performance • Latency & Cost

Making LLM inference faster and cheaper: batching, caching, quantization, and serving optimizations to cut latency and cost. Bringing enterprise performance-engineering discipline to model serving.

Inference Optimization Latency Throughput
📊

Evals & LLM Testing

Quality • Regression-Proofing

Treating LLM systems like real software: building evaluation harnesses, golden datasets, and automated scoring to catch regressions and prove quality. Turning "it feels better" into measurable, reproducible results.

Evals LLM-as-Judge Benchmarks
🚀

Building Startups

Side Projects • 0 → 1

Rapidly prototyping and shipping my own products—using AI-assisted development to go from idea to working software fast. Learning new stacks from scratch and iterating in public.

0 → 1 AI Tooling Rapid Shipping

WORK_EXPERIENCE

JPMorgan Chase & Co.

Senior Software Engineer

United States ~7 Years

🏗️ Platform & Distributed Systems

  • Designed, built, and operated production-grade backend services at enterprise scale in a highly regulated financial environment
  • Owned services end-to-end—architecture, implementation, testing, deployment, and on-call reliability
  • Collaborated across engineering, product, and security teams to ship features with strong compliance and security requirements
  • Tech Stack: Java / Spring, distributed systems, cloud, CI/CD, microservices

🤖 AI / Applied ML

  • Applying LLMs and RAG pipelines to real engineering problems—retrieval, evaluation, and serving
  • Optimizing inference latency, throughput, and cost: batching, caching, and quantization for production model serving
  • Building evaluation harnesses and LLM-as-judge pipelines to measure quality, catch regressions, and ship with confidence
  • Tech Stack: LLM APIs, RAG, Vector Databases, Python, inference serving, evals & benchmarking

TECH_STACK

AI Engineering
LLM APIs
RAG Pipelines
Vector Databases
Inference Optimization
Model Serving
Quantization
LLM Evals
LLM-as-Judge
Benchmarking
Prompt Engineering
Python
Java
Spring
Distributed Systems
Microservices
Cloud
CI/CD
System Design

WHAT_I'M_BUILDING

Inference Optimization

Ongoing

Batching • Caching • Quantization

Experimenting with serving optimizations to cut LLM latency and cost for production workloads.

Eval Harnesses

Ongoing

Golden Sets • LLM-as-Judge • Scoring

Building evaluation pipelines and benchmarks to measure LLM quality and catch regressions before they ship.

LLM App Prototypes

Ongoing

RAG • Agents • Evals

Rapidly building retrieval-augmented and agentic applications, from idea to shipped product.

Startup Projects

Ongoing

0 → 1 • Full-Stack • AI

Building my own products end-to-end, accelerated with AI-assisted development.

GET_IN_TOUCH