Software
& AI
Engineer
3+ years building large-scale distributed systems in production fintech environments, with hands-on experience architecting and shipping AI agent systems end-to-end. Directed AI coding agents (Claude Code, GitHub Copilot) to build, test, and deploy a production 6-agent LLM system, and owned high-throughput order and payment flows at 800K–1M API calls/month. Currently a Research Software Engineer at Arizona State University. MS in Data Analytics from ASU. Open to AI/LLM, Backend SWE, and Data roles.
Tarun Sai Marisetti
Backend Engineering
3+ years owning production order and payment services on a distributed transaction platform handling 800K–1M API calls/month, with fault-tolerance patterns, REST APIs, and event-driven Kafka pipelines for USAA fintech workloads.
LLM & AI Engineering
Directed AI coding agents (Claude Code, GitHub Copilot) to architect, build, test, and deploy a production 6-agent LLM system end-to-end. Built RAG pipelines with LangChain + ChromaDB, semantic chunking, query rewriting, and GPT-4o-mini reranking.
Data Analytics
MS in Data Analytics from ASU. Built predictive ML models (LightGBM, SHAP) on 911K crash records, trained a Random Forest regressor (R²=0.854) to recover missing exoplanet mass estimates, and delivered interactive data storytelling with D3.js. Skilled in end-to-end analytical pipelines from raw data to actionable insight.
- Building CI/CD pipelines via GitHub Actions for Aurora, a scientific Python library for exoplanet atmospheric retrieval, including build automation, coverage enforcement, and security scanning.
- Leveraged Claude and GitHub Copilot as daily development tools for code generation, documentation drafting, test case authoring, and codebase debugging across Aurora's scientific Python library.
- Optimizing and refactoring Aurora's core codebase, improving code quality, fixing sanity issues, and streamlining runtime performance across the retrieval pipeline.
- Integrating runtime logging and observability hooks into Aurora's retrieval pipeline to surface diagnostics across long-running scientific computing workflows.
- Authoring developer documentation on Read the Docs (API reference, installation guide) and Jupyter tutorial notebooks to support onboarding for new research contributors.
- Owned Order Service functionality on LMPS, a distributed transaction platform handling 800K–1M API calls/month across Payment, Billing, and Notification microservices deployed on Kubernetes.
- Refactored post-payment processing into a dedicated async thread pool isolated from Tomcat's HTTP pool, restoring p99 < 2s SLA compliance during billing slowness spikes.
- Diagnosed production thread pool exhaustion and JVM resource contention; cut REST timeout from 10s to 2s and tripped a Resilience4j circuit breaker to restore service stability under peak load.
- Designed and owned multi-layer idempotency on payment callbacks using application-level status checks and @Version optimistic locking, preventing duplicate billing across thousands of concurrent enterprise transactions.
- Owned migration of 5 production scheduled jobs from Ruby on Rails to Java Spring Boot (Strangler Fig), including a 6-step @Transactional cascade across 4 tables with full rollback on failure.
- Fixed a Spring AOP proxy bypass where @Async/@Transactional on same-class methods silently ran synchronously, causing incorrect post-payment execution behavior in production.
- Authored runbooks and alerting thresholds for the order service; led on-call incident response and conducted knowledge transfer sessions to onboard incoming engineers on service ownership and operational practices.
- Contributed to Spring Boot multi-service orchestration supporting USAA's financial data processing under strict regulatory audit controls, including compliance workflow coordination and ServiceNow-driven incident management.
- Investigated and resolved Tier-3 production incidents across distributed Spring services using log analysis, database inspection, and ServiceNow-driven incident workflows.
- Built unit and integration tests using JUnit and Mockito, maintaining ~80% code coverage across assigned service modules.
Cloud-Native Microservices Reliability Platform
Production-grade simulation of distributed system control-plane behavior: service discovery, auth, routing, and incident automation. Implements idempotent APIs, circuit breakers, rate limiting, and Kafka event-driven pipelines.
FinSight — Autonomous Financial Due Diligence Agent
Directed AI coding agents to architect, build, and test a production 6-agent agentic system using LangChain, functioning as a coding harness for end-to-end financial due diligence workflows over SEC 10-K/10-Q filings via the EDGAR API. Implemented LLM-driven semantic chunking, query rewriting, and GPT-4o-mini reranking (12 candidates to top 6) to maximize retrieval precision over naive vector similarity search. Validated with 34 pytest unit tests across 4 modules covering agent orchestration, reranking logic, risk score extraction, and Pydantic output contracts.
Exoplanet Habitability Atlas — ASU Research Collaborator
Built a scalable data ingestion and processing platform analyzing 4,457 exoplanets: designed an ETL pipeline with physics-based imputation, orchestrated 3 clustering algorithms, and served results via a FastAPI REST backend with a React frontend. Trained a Random Forest regressor (R²=0.854, 9 features, 300 trees, 5-fold CV) to recover mass estimates for the ~32% of planets with no direct measurement, replacing a naive power-law heuristic.
Pac-Man AI Agent — Search & Reinforcement Learning
Monte Carlo Tree Search agent for heuristic ranking and decision-making under uncertainty. Reward optimization analogous to search relevance tuning in recommendation systems.
Chicago Traffic Analytics & Crash Severity Prediction
Built a crash severity prediction pipeline on 911K records (1.66% severe class), training a LightGBM classifier achieving PR-AUC of 0.153 — outperforming Logistic Regression (0.146) and Random Forest (0.134) baselines on heavily imbalanced data. Applied F₂-based threshold tuning (threshold=0.061) to lift severe crash recall above 55%, and used SHAP analysis across the top 20 features to identify crash mechanism, temporal, and environmental drivers of severity.
Open to AI, Backend & Data Roles
Currently on F1 OPT, authorized to work in the US until Jan 2027 with STEM OPT extension available for +2 years. No employer sponsorship cost during OPT period. Based in Tempe, AZ, open to remote roles nationwide.