Software Engineer · Backend & GenAI Systems
I design scalable Python microservices and ship GenAI/LLM systems (RAG, MCP, AI agents) to production. MS CS at University of Cincinnati, 3.8 GPA, graduated May 2026.
I'm a Software Engineer with a Master of Engineering in Computer Science from the University of Cincinnati. I design, scale and deploy secure, high-availability backend systems: Python microservices, REST and gRPC APIs and cloud-native delivery, across cross-functional Agile teams.
At Capgemini I build the API gateway and microservices layer for an enterprise client, plus the GenAI on top of it: a RAG document Q&A system and a Node.js MCP server that lets LLM agents call internal APIs safely. Before that, at AgileRadioCom, I deployed PyTorch models (EKF + LSTM) into production and processed 2.2M+ records in real time; at AT&T I automated telecom network provisioning across OSS/BSS systems.
I use Claude Code as a serious engineering tool, and I care about the unglamorous parts: performance, reliability, testing and clean, maintainable code.
Real systems. Real users. Real ownership. Click How it worked on any point for the technical detail.
Multi-tenant backend transformation for an enterprise client: an API gateway and microservices aggregation layer processing high-volume customer orders, product-catalog requests and account data across legacy and cloud systems.
Downstream legacy services (billing and inventory databases) were slow and occasionally timed out at peak load. I built FastAPI microservices for the high-frequency paths and added circuit breakers (Tenacity-style) so a failing vendor or inventory API couldn't take the whole customer portal down.
async def on Uvicorn/uvloop, so the event loop yields during DB reads and external HTTP calls instead of blocking a worker.Internal service-to-service traffic over REST/JSON carried too much serialization and network overhead, so I moved the critical internal paths to gRPC with asyncio.
asyncio.gather() without blocking the main thread.The database was getting hammered by read-heavy traffic: thousands of users searching catalogs and viewing order history at once. I split the read and write paths with CQRS: writes go straight to PostgreSQL, heavy reads are served from Redis.
pool_size=20, max_overflow=10) and pool recycling, inside explicit transactions for ACID guarantees, which is what prevented connection-pool exhaustion.catalog:item:{id}. Hit → deserialized and returned in milliseconds, PostgreSQL skipped. Miss → optimized read query, then written back with a TTL.Support reps and sales-ops needed to query technical product docs and customer-policy files without digging through static PDFs. I built an internal RAG pipeline on LangChain where they ask in plain language and get grounded answers.
chunk_size=1000, overlap=200) to keep meaning across sentence boundaries.To connect the AI system to live backend data, I built an MCP server in Node.js that acts as a secure bridge: the LLM agent can call internal REST APIs (live order status, customer account tier) without any of that being hardcoded into prompts.
get_order_status(order_id), get_customer_tier(account_id)) has a strict JSON schema; arguments are validated before any request runs, so malformed calls never reach the backend.Every service and AI workload ships through automated pipelines onto Kubernetes, with rolling updates so a release never interrupts live users.
kubectl apply from CI.maxSurge: 25%, maxUnavailable: 0; readiness probes on /healthz so traffic only reaches warmed-up pods, liveness probes restart unresponsive ones.Backend operations platform that automates network provisioning and service activation across OSS/BSS systems: the bridge between a customer placing an order and the network infrastructure actually configuring and activating that service.
Originally, operators manually verified network capacity and pushed configurations across legacy databases for every enterprise provisioning request. I built microservices that ingest those orders, validate inventory and trigger network-activation APIs automatically.
202 Accepted with an order_id, handing heavy orchestration to a background queue (Celery / Redis / RabbitMQ) so HTTP calls never time out at peak.PENDING → VALIDATING → PROVISIONING → ACTIVE.tenacity absorb transient network failures without operator intervention.Activation pipelines touch critical infrastructure, so a production bug could take down live customer links. I built PyTest suites for API/unit tests and Selenium for end-to-end regression, wired into Jenkins with quality gates that block a merge or deploy if they aren't met.
unittest.mock / pytest-mock) so tests never touch live telecom hardware.@pytest.mark.parametrize pushes hundreds of malformed payloads (bad VLAN IDs, missing bandwidth, expired tokens) through the same test.selenium/standalone-chrome) nightly and on release builds.pytest --cov and fails the build below 85% coverage or on any high-severity SonarQube finding, posting status checks that block pull requests.When provisioning failed (port timeouts, invalid VLAN tags, legacy API timeouts), support engineers spent hours hunting through raw logs scattered across systems. I helped build an internal GenAI log-analysis service that finds the root cause and writes plain-English fixes.
order_id, device IP, error code, timestamp) and stitches logs from microservices and legacy devices into one incident trace via the order_id correlation ID.The database held millions of rows of inventory and order state, and at peak provisioning hours long-running queries bottlenecked the API. I traced the slow queries, fixed the SQL and indexes, and reshaped how Python handled the data.
pg_stat_statements plus log_min_duration_statement (queries over ~200 ms) to find the offenders; EXPLAIN (ANALYZE, BUFFERS) to spot sequential scans, inefficient nested-loop joins and heavy disk I/O.(status, location_id)) turned full scans into index scans.INNER JOINs so the planner can use hash and merge joins.joinedload() / selectinload(); replaced SELECT * with explicit field selection and added cursor-based pagination to cap memory and payload size.Hover a preview (tap on mobile) to see it run.
file:line citations and a call-graph, running 100% locally, no API keys. I rebuilt a naive text-chunk RAG (citation accuracy 32.7%) into an AST-aware engine: tree-sitter chunks, a symbol/call graph in SQLite, BM25 + dense retrieval fused with RRF and cross-encoder re-ranking. Citation accuracy is now 100% on its benchmark.to_tsvector) with a debounced search UI, end-to-end typed via tRPC and covered by Jest tests.coverage.xml to find untested functions, builds targeted prompts, calls the OpenAI API to write pytest tests, then re-runs pytest --cov to measure the improvement, a feedback loop between coverage data and generative AI.The tools I work with every day.
Always up for a conversation about backend engineering, GenAI systems and platform work.