Ayush Gupta

resume

Noida, India.

Staff Software Engineer & Founding Engineer specializing in high-scale real-time platforms, production-grade GenAI systems, Advanced RAG, Agentic workflows, and distributed AI inference. Spearheaded the Vooz ecosystem serving 500k+ MAU, engineering Triton & vLLM inference clusters (1280+ RPS), OpenTelemetry observability, and zero-latency browser-edge ONNX/WebGPU safety layers across multi-cloud environments (Azure AKS, ArgoCD GitOps, DigitalOcean).

"Simplicity is a prerequisite for reliability."

— Edsger W. Dijkstra

skills & stack

S

Core technologies, architecture paradigms, and domain skills.

Frontend

Next.js 16 (Parallel/Intercepting Routes), TypeScript, Redux Toolkit, React Native CLI / Expo (Android/iOS), WebRTC, Framer Motion

Backend

Python, FastAPI, Node.js, Express, .NET, Golang, gRPC, CQRS, Event-Driven

AI Engineering

LLM APIs (Anthropic Claude, OpenAI), RAG, Triton Inference Server, vLLM, ONNX Runtime, LangChain, Prompt Engineering, Model Quantization

DevOps & Cloud

Azure AKS, ArgoCD (GitOps), Azure Front Door (AFD), AWS, OpenTelemetry (OTel), Prometheus, Grafana

Data & Storage

Memgraph (Graph DB), PostgreSQL, Redis, Cache API (Service Worker), Persistence Layer

Real-Time & Web3

SignalR Hub Sync, Socket.io, P2P Data Channels, MediaStream Manipulation, TURN/STUN, Solana (web3.js), SPL Tokens, Helius Webhooks, Phantom Wallet

Leadership & Governance

Founding Team Leadership, Technical Mentorship, Architectural Governance, Product Strategy

experience

E

Organizations and ventures where I've architected production infrastructure and led core engineering teams.

Staff Software Engineer (Founding Engineer) @ VOOZ INC

Sep 2024 — Present
  • Architectural Foundation: Engineered multi-tenant Next.js environment utilizing Parallel & Intercepting Routes to deliver seamless navigation scaling to 500k+ MAU with optimized SSR/caching delivery.
  • Real-Time Video Coordination: Designed and maintained the Vooz Video Coordinator state machine managing WebRTC handshakes and Media Recovery Logic for hot-swapping hardware mid-call.
  • AI Inference at Scale: Spearheaded migration to high-throughput Triton Inference Server & vLLM cluster with Dynamic Batching achieving 1280+ RPS and full OpenTelemetry (OTel) tracing.
  • Web3 Digital Economy: Architected end-to-end Solana payment infrastructure Node.js microservice via Helius Webhooks for atomic transaction confirmation and point-balance hydration.
  • LLM Orchestration: Integrated LLM APIs (Anthropic Claude, OpenAI) to power conversational features, automated moderation, and semantic search queries.
  • Infrastructure & Edge Logic: Orchestrated multi-cloud deployments (Azure AKS / DigitalOcean) secured by ArgoCD (GitOps) and Azure Front Door (AFD) with cookie-based edge routing.
  • System Resilience: Established zero-latency SignalR Hub Synchronization protocol to recover user state from abnormal disconnections (Code 1006) for 99.9% session continuity.
  • Client-Side Intelligence & Safety: Engineered browser-local inference engine using ONNX Runtime (WebGPU/WASM) and Cache API for real-time zero-latency NSFW content filtering.

Software Developer (Lead Role) @ AIMICA LTD

Jan 2023 — Sep 2024
  • Led a development team of 3 building AI-powered mobile applications using React Native (Expo).
  • Designed & deployed scalable RAG-based knowledge system integrating Anthropic Claude, OpenAI GPT-4, and open-source LLMs.
  • Developed backend orchestration services in Node.js & Python for semantic retrieval workflows and LLM caching.
  • Authored custom native modules in Swift & Java bridging native mobile features with AI context layers.

Mobile Developer @ INFOEDGE SOFTWARE SOLUTIONS

Dec 2021 — Dec 2022
  • Optimized critical React Native application flows for high traffic, eliminating memory leaks and improving frame rates.
  • Collaborated with backend teams to strictly type REST & GraphQL API contracts for cross-platform data consistency.

work & systems

P

Featured architectures, production platforms, and AI engineering projects.

Triton & vLLM Distributed AI Inference Cluster

Triton / vLLM / OTel

High-throughput GPU inference backend powering 1280+ RPS across production LLM and vision workloads. Engineered with Dynamic Batching, Prometheus metrics, and OpenTelemetry distributed tracing.

Zero-Latency Edge ONNX Safety Engine

ONNX / WebGPU / WASM

Browser-native inference engine evaluating real-time video feeds on-device using WebGPU execution providers and Cache API, eliminating cloud inference overhead.

Medical Paper NLP Categoriser

Python / TensorFlow / Pandas

Trained deep learning models using TensorFlow on PubMed RCT datasets. Achieved 90%+ accuracy in classifying unstructured medical research abstracts.

Secure WebRTC Video Coordinator

WebRTC / SignalR / Socket.io

Cross-platform signaling and WebRTC state machine featuring end-to-end encryption, Media Recovery logic, and SignalR hub sync for 99.9% session continuity.

education

B.Tech in Computer Science Engineering

Dr. A.P.J. Abdul Kalam Technical University

2018 — 2022