Ayush Gupta

resume
Available for contract work — AI inference, WebRTC & real-time systems what I do →

Staff software engineer and founding engineer. For the past two years I've built the infrastructure underneath a real-time video platform serving 500k+ monthly active users — the cluster that serves its models, the WebRTC coordinator that keeps calls alive through bad networks, and the multi-cloud deployment underneath both.

I work where AI systems meet real-time constraints: GPU serving that has to be fast and cheap at the same time, video that isn't allowed to drop, and retrieval pipelines that survive contact with production. Triton and vLLM, ONNX on WebGPU, Next.js and React Native, Python and Go, Kubernetes and ArgoCD — with OpenTelemetry through all of it, because systems you can't see into are systems you can't fix.

500k+
monthly active users served
1,280+
inference requests / second
99.9%
WebRTC session continuity
5+ yrs
shipping production systems

skills & stack

S

Core technologies, architecture paradigms, and domain skills.

Frontend

Next.js 16 (Parallel/Intercepting Routes), TypeScript, Redux Toolkit, React Native CLI / Expo (Android/iOS), WebRTC, Framer Motion

Backend

Python, FastAPI, Node.js, Express, .NET, Golang, gRPC, CQRS, Event-Driven

AI Engineering

LLM APIs (Anthropic Claude, OpenAI), RAG, Triton Inference Server, vLLM, ONNX Runtime, LangChain, LangGraph, AutoGen, Prompt Engineering, Model Quantization

DevOps & Cloud

Azure AKS, ArgoCD (GitOps), Azure Front Door (AFD), AWS, OpenTelemetry (OTel), Prometheus, Grafana

Data & Storage

Memgraph (Graph DB), PostgreSQL, Redis, Cache API (Service Worker), Persistence Layer

Real-Time & Web3

SignalR Hub Sync, Socket.io, P2P Data Channels, MediaStream Manipulation, TURN/STUN, Solana (web3.js), SPL Tokens, Helius Webhooks, Phantom Wallet

Leadership & Governance

Founding Team Leadership, Technical Mentorship, Architectural Governance, Product Strategy

experience

E

Organizations and ventures where I've architected production infrastructure and led core engineering teams.

Staff Software Engineer (Founding Engineer) @ VOOZ INC

Sep 2024 — Present
  • Architectural Foundation: Engineered multi-tenant Next.js environment utilizing Parallel & Intercepting Routes to deliver seamless navigation scaling to 500k+ MAU with optimized SSR/caching delivery.
  • Real-Time Video Coordination & Resilience: Designed and maintained the Vooz Video Coordinator state machine managing WebRTC handshakes and Media Recovery Logic for hot-swapping hardware mid-call; established a SignalR Hub Synchronization protocol recovering user state from abnormal disconnections (Code 1006) for 99.9% session continuity.
  • Group Audio Platform (0→1): Shipped Hangouts, host-owned group audio rooms on a P2P WebRTC mesh, holding all live room state in Redis behind nine atomic Lua scripts; time-scored membership makes crashed clients and half-open sockets self-expire with no reconciliation job.
  • Service Decomposition: Split a monolithic .NET API into independently scaled REST and Realtime services behind one ingress, decoupling WebSocket connection density from REST capacity so hub headroom no longer forces redundant REST pods.
  • Database Performance: Resolved production PostgreSQL CPU saturation (~97% → nominal) by rewriting the six hottest queries by pg_stat_statements share — OR→UNION decomposition, covering indexes built CONCURRENTLY, and a username cache eliminating 443M redundant lookups.
  • AI Inference at Scale: Spearheaded migration to high-throughput Triton Inference Server & vLLM cluster with Dynamic Batching achieving 1280+ RPS and full OpenTelemetry (OTel) tracing.
  • Multi-Rail Payment Infrastructure: Architected payments across crypto and fiat — Solana (Node.js microservice, Helius Webhooks, atomic confirmation) alongside CCBill and Square (Web Payments SDK, 3-D Secure buyer verification, Apple Pay) behind one provider-switchable gateway.
  • LLM Orchestration: Integrated LLM APIs (Anthropic Claude, OpenAI) to power conversational features, automated moderation, and semantic search queries.
  • Infrastructure & Edge Logic: Orchestrated multi-cloud deployments (Azure AKS / DigitalOcean) secured by ArgoCD (GitOps) and Azure Front Door (AFD) with cookie-based edge routing.
  • Client-Side Intelligence & AI Safety: Engineered browser-local inference engine (ONNX Runtime, WebGPU/WASM in a Web Worker) scanning remote video feeds in real time, delivered via the Cache API; paired with backend NSFW services in a dual-track pipeline for a unified blur/report response.

Software Developer (Lead Role) @ AIMICA LTD

Jan 2023 — Sep 2024
  • Led a development team of 3 building AI-powered mobile applications using React Native (Expo).
  • Designed & deployed scalable RAG & Multimodal AI systems integrating Claude, OpenAI GPT, LLaMA, image generation, and open-source AI models.
  • Developed backend orchestration services in Node.js & Python for semantic retrieval workflows and LLM caching.
  • Authored custom native modules in Swift & Java bridging native mobile features with AI context layers.

Mobile Developer @ INFOEDGE SOFTWARE SOLUTIONS

Dec 2021 — Dec 2022
  • Optimized critical React Native application flows for high traffic, eliminating memory leaks and improving frame rates.
  • Collaborated with backend teams to strictly type REST & GraphQL API contracts for cross-platform data consistency.

work & systems

P

Featured architectures, production platforms, and AI engineering projects.

Triton & vLLM Distributed AI Inference Cluster

Triton / vLLM / OTel

High-throughput GPU inference backend powering 1280+ RPS across production LLM and vision workloads. Engineered with Dynamic Batching, Prometheus metrics, and OpenTelemetry distributed tracing.

read the write-up

Zero-Latency Edge ONNX Safety Engine

ONNX / WebGPU / WASM

Browser-native inference engine evaluating real-time video feeds on-device using WebGPU execution providers and Cache API, eliminating cloud inference overhead.

Medical Paper NLP Categoriser

Python / TensorFlow / Pandas

Trained deep learning models using TensorFlow on PubMed RCT datasets. Achieved 90%+ accuracy in classifying unstructured medical research abstracts.

Secure WebRTC Video Coordinator

WebRTC / SignalR / Socket.io

Cross-platform signaling and WebRTC state machine featuring end-to-end encryption, Media Recovery logic, and SignalR hub sync for 99.9% session continuity.

education

B.Tech in Computer Science Engineering

Dr. A.P.J. Abdul Kalam Technical University

2018 — 2022