Open to opportunities

Varun
Kumar
Andigari

Software Engineer · Backend & GenAI Systems

I design scalable Python microservices and ship GenAI/LLM systems (RAG, MCP, AI agents) to production. MS CS at University of Cincinnati, 3.8 GPA, graduated May 2026.

3.8
GPA / 4.0
2.2M+
Records in prod
3+
Years exp
request_dodger.py
score 0 ♥♥♥ best 0
200 OK = points · 500 / timeout = lose a life · CACHE = shield
mouse / touch / ← → to move
Python  •  FastAPI  •  Microservices  •  gRPC  •  RAG Systems  •  LLM / GenAI  •  MCP  •  Kubernetes  •  PostgreSQL  •  Redis  •  AWS  •  CI/CD  •  Python  •  FastAPI  •  Microservices  •  gRPC  •  RAG Systems  •  LLM / GenAI  •  MCP  •  Kubernetes  •  PostgreSQL  •  Redis  •  AWS  •  CI/CD  • 
About me

Builder at the intersection
of software & AI.

I'm a Software Engineer with a Master of Engineering in Computer Science from the University of Cincinnati. I design, scale and deploy secure, high-availability backend systems: Python microservices, REST and gRPC APIs and cloud-native delivery, across cross-functional Agile teams.

At Capgemini I build the API gateway and microservices layer for an enterprise client, plus the GenAI on top of it: a RAG document Q&A system and a Node.js MCP server that lets LLM agents call internal APIs safely. Before that, at AgileRadioCom, I deployed PyTorch models (EKF + LSTM) into production and processed 2.2M+ records in real time; at AT&T I automated telecom network provisioning across OSS/BSS systems.

I use Claude Code as a serious engineering tool, and I care about the unglamorous parts: performance, reliability, testing and clean, maintainable code.

3.8
GPA / 4.0
3+
Years building
2.2M+
Records processed in production
Experience

Where I've worked.

Real systems. Real users. Real ownership. Click How it worked on any point for the technical detail.

Capgemini
Jan 2026 to Present
Remote, USA
Software Engineer

Multi-tenant backend transformation for an enterprise client: an API gateway and microservices aggregation layer processing high-volume customer orders, product-catalog requests and account data across legacy and cloud systems.

  • Downstream legacy services (billing and inventory databases) were slow and occasionally timed out at peak load. I built FastAPI microservices for the high-frequency paths and added circuit breakers (Tenacity-style) so a failing vendor or inventory API couldn't take the whole customer portal down.

    • Async handlers: endpoints are async def on Uvicorn/uvloop, so the event loop yields during DB reads and external HTTP calls instead of blocking a worker.
    • Process model: Gunicorn managing Uvicorn workers, one per CPU core, to use the full machine.
    • Validation: Pydantic for fast request parsing and serialization, keeping CPU overhead per request low.
    • Connection pooling: persistent pools for PostgreSQL (SQLAlchemy async sessions) and Redis, removing per-request TCP handshake latency.
    Caching + async pipeline: avg API latency down 30%
  • Internal service-to-service traffic over REST/JSON carried too much serialization and network overhead, so I moved the critical internal paths to gRPC with asyncio.

    • Protobuf: compact binary serialization instead of text JSON, cutting payloads by roughly 60 to 70% on the wire.
    • HTTP/2 multiplexing: many concurrent requests over one long-lived connection, with no head-of-line blocking, no repeated connection setup.
    • asyncio fan-out: with gRPC's async stubs, a service fires dozens of inter-service calls at once via asyncio.gather() without blocking the main thread.
  • The database was getting hammered by read-heavy traffic: thousands of users searching catalogs and viewing order history at once. I split the read and write paths with CQRS: writes go straight to PostgreSQL, heavy reads are served from Redis.

    • Command path (writes): SQLAlchemy AsyncSession over asyncpg with explicit pool bounds (pool_size=20, max_overflow=10) and pool recycling, inside explicit transactions for ACID guarantees, which is what prevented connection-pool exhaustion.
    • Query path (reads): cache-aside on Redis with structured keys like catalog:item:{id}. Hit → deserialized and returned in milliseconds, PostgreSQL skipped. Miss → optimized read query, then written back with a TTL.
    • Consistency: every write invalidates or updates the affected Redis key, with a 5 to 15 minute TTL as a safety net against stale data.
  • Support reps and sales-ops needed to query technical product docs and customer-policy files without digging through static PDFs. I built an internal RAG pipeline on LangChain where they ask in plain language and get grounded answers.

    • Ingestion: PDFs, policy files and manuals parsed and split with a recursive character splitter (chunk_size=1000, overlap=200) to keep meaning across sentence boundaries.
    • Vector store: chunks embedded and stored in a vector database (PostgreSQL/pgvector or Redis vector search).
    • Retrieval: embed the question, run cosine similarity, take the top 3 to 5 chunks and inject them into the prompt as context.
    • Hallucination control: the system prompt says to answer only from the provided context, and to say "I don't know" when it isn't there.
  • To connect the AI system to live backend data, I built an MCP server in Node.js that acts as a secure bridge: the LLM agent can call internal REST APIs (live order status, customer account tier) without any of that being hardcoded into prompts.

    • Tool schemas: each tool (e.g. get_order_status(order_id), get_customer_tier(account_id)) has a strict JSON schema; arguments are validated before any request runs, so malformed calls never reach the backend.
    • Auth: every request carries a scoped OAuth 2.0 JWT for the agent's session; the server verifies the signature and enforces RBAC so a user can't read data above their permission level.
  • Every service and AI workload ships through automated pipelines onto Kubernetes, with rolling updates so a release never interrupts live users.

    • CI (Jenkins): on push, runs lint, static analysis and unit/integration tests (PyTest/Jest); if gates pass, builds a multi-stage Docker image tagged with the Git commit SHA and pushes it to the registry.
    • CD (ArgoCD, GitOps): Jenkins updates the image tag in a manifest repo; ArgoCD detects the drift and syncs the cluster, with no imperative kubectl apply from CI.
    • Zero downtime: RollingUpdate with maxSurge: 25%, maxUnavailable: 0; readiness probes on /healthz so traffic only reaches warmed-up pods, liveness probes restart unresponsive ones.
    • Security: OAuth 2.0 + JWT at the gateway, signatures checked against JWKS, roles and scopes enforced (RBAC) before a request reaches any internal service.
    Deployment cycles down 60%, zero downtime
AgileRadioCom LLC
May 2025 to Dec 2025
Remote, USA
AI Software Engineer Intern
  • Deployed PyTorch ML models (EKF, online LSTM) into production for real-time state estimation, forecasting and autonomous tuning, integrating GenAI/LLMs to interpret model outputs.
  • Engineered a Python/Flask AI backend with Redis caching and WebSocket dashboards (Chart.js), streaming multi-API data to process 2.2M+ records and render real-time predictions, accuracy metrics and LLM insights.
  • Resolved a 146 GB log-growth production outage, freeing 157 GB of disk space through root-cause analysis, while refactoring a 1,200-line monolithic codebase into a modular Python architecture for scalable AI/ML workflows.
AT&T
Dec 2022 to Jul 2024
Hyderabad, India
Software Engineer

Backend operations platform that automates network provisioning and service activation across OSS/BSS systems: the bridge between a customer placing an order and the network infrastructure actually configuring and activating that service.

  • Originally, operators manually verified network capacity and pushed configurations across legacy databases for every enterprise provisioning request. I built microservices that ingest those orders, validate inventory and trigger network-activation APIs automatically.

    • Async ingestion: orders arrive via REST webhooks; the API validates the payload with Pydantic and immediately returns 202 Accepted with an order_id, handing heavy orchestration to a background queue (Celery / Redis / RabbitMQ) so HTTP calls never time out at peak.
    • Inventory validation: workers query PostgreSQL network inventory for port availability, circuit paths and bandwidth constraints.
    • No double-booking: Redis distributed locks prevent two orders from claiming the same port; each order moves through an explicit state machine: PENDING → VALIDATING → PROVISIONING → ACTIVE.
    • Legacy adapters: lightweight Node.js/Flask wrappers translate modern JSON into the legacy SOAP/XML that downstream network controllers expect.
    • Resilience: exponential backoff and retry policies with Python's tenacity absorb transient network failures without operator intervention.
    Data-processing efficiency +30%
  • Activation pipelines touch critical infrastructure, so a production bug could take down live customer links. I built PyTest suites for API/unit tests and Selenium for end-to-end regression, wired into Jenkins with quality gates that block a merge or deploy if they aren't met.

    • Isolation: external OSS/BSS responses and third-party webhooks are mocked (unittest.mock / pytest-mock) so tests never touch live telecom hardware.
    • Fixtures: reusable PyTest fixtures spin up isolated PostgreSQL databases in Docker, run inside rollback-able transactions and tear down automatically, with no cross-test contamination.
    • Edge cases at scale: @pytest.mark.parametrize pushes hundreds of malformed payloads (bad VLAN IDs, missing bandwidth, expired tokens) through the same test.
    • E2E: Selenium WebDriver against the provisioning-agent portal using the Page Object Model, so UI tweaks don't break the suite; run headless in Docker (selenium/standalone-chrome) nightly and on release builds.
    • Gates: Jenkins runs pytest --cov and fails the build below 85% coverage or on any high-severity SonarQube finding, posting status checks that block pull requests.
    Coverage +50% · regression cycles down 30%
  • When provisioning failed (port timeouts, invalid VLAN tags, legacy API timeouts), support engineers spent hours hunting through raw logs scattered across systems. I helped build an internal GenAI log-analysis service that finds the root cause and writes plain-English fixes.

    • Correlation: a Python pipeline ingests failure events, strips heartbeat noise, extracts identifiers (order_id, device IP, error code, timestamp) and stitches logs from microservices and legacy devices into one incident trace via the order_id correlation ID.
    • Token control: a sliding window of about five minutes around the failure isolates stack traces, 4xx/5xx statuses and hardware-timeout signals, keeping context small and API calls fast and cheap.
    • Structured prompting: role context + incident telemetry + an enforced JSON output: root cause, confidence score, affected subsystem and recommended steps.
    • Grounding: the model may only diagnose from the supplied logs and known operational playbooks, preventing invented hardware commands.
    • Delivery: summaries land on support dashboards (or Slack/Teams webhooks) so engineers can review and act in minutes.
    Lower mean time to resolution for provisioning incidents
  • Automated Python test execution and deployment via Jenkins CI/CD pipelines with automated release gates, reducing manual effort by 40% and ensuring Dev/QA environment consistency.
  • The database held millions of rows of inventory and order state, and at peak provisioning hours long-running queries bottlenecked the API. I traced the slow queries, fixed the SQL and indexes, and reshaped how Python handled the data.

    • Diagnosis: pg_stat_statements plus log_min_duration_statement (queries over ~200 ms) to find the offenders; EXPLAIN (ANALYZE, BUFFERS) to spot sequential scans, inefficient nested-loop joins and heavy disk I/O.
    • Indexing: composite indexes on co-queried columns (e.g. (status, location_id)) turned full scans into index scans.
    • Query refactor: unindexed subqueries and outer joins rewritten as explicit INNER JOINs so the planner can use hash and merge joins.
    • ORM fixes: eliminated N+1 queries with joinedload() / selectinload(); replaced SELECT * with explicit field selection and added cursor-based pagination to cap memory and payload size.
    Avg API response time down 25%
Projects

What I've built.

Hover a preview (tap on mobile) to see it run.

01 · AI Agent
Codebase QA
Ask a natural-language question about any GitHub repo and get an answer with real file:line citations and a call-graph, running 100% locally, no API keys. I rebuilt a naive text-chunk RAG (citation accuracy 32.7%) into an AST-aware engine: tree-sitter chunks, a symbol/call graph in SQLite, BM25 + dense retrieval fused with RRF and cross-encoder re-ranking. Citation accuracy is now 100% on its benchmark.
Python tree-sitter ChromaDB BM25 + RRF SQLite Ollama Streamlit
View on GitHub →
# hybrid retrieval → cited answer
def ask(question):
hits = rrf(bm25(q), dense(q))
top = rerank(hits)[:5]
calls = expand_call_graph(top)
return answer(top, calls)
# Citation accuracy: 32.7% → 100%
localhost:8501 · Codebase QA
Ask about the repo
BM25denseRRFre-rank
authenticate() login.py:31-58 hashes the input with verify_password hashing.py:22-40 and compares it to the record from get_user users.py:12-19.
authenticate verify_password get_user
Hover to previewTap to previewSimulated preview
02 · Full Stack
Link Vault
A minimalist bookmark manager. Paste a URL and it scrapes the page's title, description and preview image, auto-tags it by domain, and upserts it into Postgres so re-adding a link refreshes rather than duplicates. Every saved link is full-text searchable (to_tsvector) with a debounced search UI, end-to-end typed via tRPC and covered by Jest tests.
Next.js TypeScript tRPC Prisma PostgreSQL Tailwind Jest
View on GitHub →
// tRPC mutation
addBookmark: procedure
.input(z.object({ url }))
.mutation(async ({ input }) => {
const meta = await scrape(input.url)
return db.upsert(meta, autoTag(meta))
})
// full-text search: tsvector
localhost:3000 · Link Vault
Add
G
varunk61/codebase-qagithub.com · metadata fetcheddevcode
search
N
Next.js docsnextjs.org
t
tRPC: end-to-end typesafe APIstrpc.io
Hover to previewTap to previewSimulated preview
03 · Agentic AI
MCP Enterprise Integration Hub
A governed AI-agent platform that exposes GitHub Issues, Jira and Slack as MCP servers. A LangGraph agent routes plain-language requests through Pydantic-validated tool schemas, with a mandatory human-in-the-loop approval gate before any write. Every decision, tool call and outcome is persisted to PostgreSQL for a full audit trail, and PGVector surfaces similar past actions before each new one is planned.
Python FastAPI LangGraph MCP PGVector PostgreSQL Pydantic Docker
View on GitHub →
# LangGraph flow
parse_intent → retrieve_similar # pgvector
→ plan_action → hitl_gate
if plan.is_write:
await human_approval()
→ execute("github" | "jira" | "slack")
→ log_run() # full audit trail
POST /agent/run
>
✓ parse_intent → jira.create_issue
✓ retrieve_similar · PGVector · past actions found
✓ plan_action · params validated (Pydantic)
⏸ hitl_gate · write op → pending_approval
Approve
✓ execute · jira.create_issue → ABC-123
✓ audit log → PostgreSQL
Hover to previewTap to previewSimulated preview
04 · Deep Learning
Brain Tumor Detection
Transfer-learning classifier that reads an MRI scan and predicts tumor / no tumor, built on an ImageNet-pretrained EfficientNetB0 with a custom head (~95% validation accuracy). Served through a Flask + Gunicorn web app with upload validation, and a Grad-CAM heatmap that highlights the region the model focused on, so predictions are interpretable, not a black box.
TensorFlow / Keras EfficientNetB0 Grad-CAM Flask Gunicorn OpenCV
View on GitHub →
base = EfficientNetB0(weights="imagenet")
base.trainable = False
x = GlobalAveragePooling2D()(base.output)
x = Dense(64, "relu")(x)
x = Dropout(0.5)(x)
out = Dense(2, "softmax")(x)
# ~95% val accuracy · Grad-CAM explainability
localhost:5000 · MRI classifier
↑ scan.jpg uploaded
preprocess → 224×224×3
EfficientNetB0 · inference
Prediction: tumor
confidence
✓ Grad-CAM: focus region
Hover to previewTap to previewSimulated preview · illustrative
05 · GenAI · Course project (team)
Automated Unit Test Generation
Generates unit tests with an LLM to close coverage gaps. It parses a Cobertura coverage.xml to find untested functions, builds targeted prompts, calls the OpenAI API to write pytest tests, then re-runs pytest --cov to measure the improvement, a feedback loop between coverage data and generative AI.
Python OpenAI API Prompt Engineering pytest pytest-cov
View on GitHub →
# coverage-driven test generation
gaps = parse_coverage("coverage.xml")
for fn in gaps.untested:
prompt = build_prompt(fn)
tests = openai_generate(prompt)
save(f"test_{fn.name}.py", tests)
# then: pytest --cov → re-measure
terminal
$
coverage%
$
✓ test_add.py
✓ test_subtract.py  ✓ test_multiply.py
✓ test_divide.py  ✓ test_modulus.py
$
coverage%
Hover to previewTap to previewSimulated preview · illustrative
Skills

What I reach for.

The tools I work with every day.

Backend & Programming
PythonJavaScriptTypeScriptFastAPIFlaskDjangoAsyncIOMultithreadingMultiprocessingSQLAlchemyPydanticPytestScripting & AutomationOOPData Structures & AlgorithmsPoetryPip
AI, Generative AI & LLM
Large Language ModelsOpenAI APIPrompt EngineeringRAGVector SearchSemantic SearchLangChainAI AgentsFunction CallingModel Context Protocol (MCP)Context Management
Machine Learning & Data
Supervised & Unsupervised LearningScikit-learnPandasNumPyData PreprocessingFeature EngineeringModel EvaluationData AnalysisData VisualizationJupyter Notebook
Microservices & API Development
RESTful Web ServicesREST APIsAPI GatewayOpenAPI/SwaggerJSONXMLService DiscoveryDistributed SystemsEvent-Driven ArchitectureOAuth 2.0JWTAPI Authentication & Authorization
Cloud & DevOps
AWS (EC2, S3, RDS, Lambda, CloudWatch)DockerKubernetesJenkinsCI/CDTerraformContainerizationInfrastructure as Code (IaC)Cloud DeploymentMonitoring & Logging
Data Infrastructure, Messaging & Testing
PostgreSQLMySQLMongoDBRedisSQLData ModelingQuery OptimizationApache KafkaRabbitMQCeleryAsync ProcessingPytestAPI TestingPostmanTDD
Development Tools & Engineering
GitJiraConfluenceLinuxAgile/ScrumSystem DesignClaude Code (Agentic Workflows)Debugging & Troubleshooting
Contact
Let's build
something
great.

Always up for a conversation about backend engineering, GenAI systems and platform work.