Ayush Gupta
resumeNoida, India.
Staff Software Engineer & Founding Engineer specializing in high-scale real-time platforms, production-grade GenAI systems, Advanced RAG, Agentic workflows, and distributed AI inference. Spearheaded the Vooz ecosystem serving 500k+ MAU, engineering Triton & vLLM inference clusters (1280+ RPS), OpenTelemetry observability, and zero-latency browser-edge ONNX/WebGPU safety layers across multi-cloud environments (Azure AKS, ArgoCD GitOps, DigitalOcean).
"Simplicity is a prerequisite for reliability."
— Edsger W. Dijkstra
skills & stack
SCore technologies, architecture paradigms, and domain skills.
Frontend
Next.js 16 (Parallel/Intercepting Routes), TypeScript, Redux Toolkit, React Native CLI / Expo (Android/iOS), WebRTC, Framer Motion
Backend
Python, FastAPI, Node.js, Express, .NET, Golang, gRPC, CQRS, Event-Driven
AI Engineering
LLM APIs (Anthropic Claude, OpenAI), RAG, Triton Inference Server, vLLM, ONNX Runtime, LangChain, Prompt Engineering, Model Quantization
DevOps & Cloud
Azure AKS, ArgoCD (GitOps), Azure Front Door (AFD), AWS, OpenTelemetry (OTel), Prometheus, Grafana
Data & Storage
Memgraph (Graph DB), PostgreSQL, Redis, Cache API (Service Worker), Persistence Layer
Real-Time & Web3
SignalR Hub Sync, Socket.io, P2P Data Channels, MediaStream Manipulation, TURN/STUN, Solana (web3.js), SPL Tokens, Helius Webhooks, Phantom Wallet
Leadership & Governance
Founding Team Leadership, Technical Mentorship, Architectural Governance, Product Strategy
experience
EOrganizations and ventures where I've architected production infrastructure and led core engineering teams.
Staff Software Engineer (Founding Engineer) @ VOOZ INC
Sep 2024 — Present- Architectural Foundation: Engineered multi-tenant Next.js environment utilizing Parallel & Intercepting Routes to deliver seamless navigation scaling to 500k+ MAU with optimized SSR/caching delivery.
- Real-Time Video Coordination: Designed and maintained the Vooz Video Coordinator state machine managing WebRTC handshakes and Media Recovery Logic for hot-swapping hardware mid-call.
- AI Inference at Scale: Spearheaded migration to high-throughput Triton Inference Server & vLLM cluster with Dynamic Batching achieving 1280+ RPS and full OpenTelemetry (OTel) tracing.
- Web3 Digital Economy: Architected end-to-end Solana payment infrastructure Node.js microservice via Helius Webhooks for atomic transaction confirmation and point-balance hydration.
- LLM Orchestration: Integrated LLM APIs (Anthropic Claude, OpenAI) to power conversational features, automated moderation, and semantic search queries.
- Infrastructure & Edge Logic: Orchestrated multi-cloud deployments (Azure AKS / DigitalOcean) secured by ArgoCD (GitOps) and Azure Front Door (AFD) with cookie-based edge routing.
- System Resilience: Established zero-latency SignalR Hub Synchronization protocol to recover user state from abnormal disconnections (Code 1006) for 99.9% session continuity.
- Client-Side Intelligence & Safety: Engineered browser-local inference engine using ONNX Runtime (WebGPU/WASM) and Cache API for real-time zero-latency NSFW content filtering.
Software Developer (Lead Role) @ AIMICA LTD
Jan 2023 — Sep 2024- Led a development team of 3 building AI-powered mobile applications using React Native (Expo).
- Designed & deployed scalable RAG-based knowledge system integrating Anthropic Claude, OpenAI GPT-4, and open-source LLMs.
- Developed backend orchestration services in Node.js & Python for semantic retrieval workflows and LLM caching.
- Authored custom native modules in Swift & Java bridging native mobile features with AI context layers.
Mobile Developer @ INFOEDGE SOFTWARE SOLUTIONS
Dec 2021 — Dec 2022- Optimized critical React Native application flows for high traffic, eliminating memory leaks and improving frame rates.
- Collaborated with backend teams to strictly type REST & GraphQL API contracts for cross-platform data consistency.
work & systems
PFeatured architectures, production platforms, and AI engineering projects.
Triton & vLLM Distributed AI Inference Cluster
Triton / vLLM / OTelHigh-throughput GPU inference backend powering 1280+ RPS across production LLM and vision workloads. Engineered with Dynamic Batching, Prometheus metrics, and OpenTelemetry distributed tracing.
Zero-Latency Edge ONNX Safety Engine
ONNX / WebGPU / WASMBrowser-native inference engine evaluating real-time video feeds on-device using WebGPU execution providers and Cache API, eliminating cloud inference overhead.
Medical Paper NLP Categoriser
Python / TensorFlow / PandasTrained deep learning models using TensorFlow on PubMed RCT datasets. Achieved 90%+ accuracy in classifying unstructured medical research abstracts.
Secure WebRTC Video Coordinator
WebRTC / SignalR / Socket.ioCross-platform signaling and WebRTC state machine featuring end-to-end encryption, Media Recovery logic, and SignalR hub sync for 99.9% session continuity.
education
B.Tech in Computer Science Engineering
Dr. A.P.J. Abdul Kalam Technical University