Sahil Tomar (dev-S-t)

Title: AI Solutions Engineer @ TechieMaya

Handles: GitHub: dev-S-t | LinkedIn: dev-s-t

Location: Ghaziabad, Uttar Pradesh, India (Open to Remote / Delhi NCR / Bengaluru / Pune)

About Sahil Tomar (dev-S-t)

Sahil Tomar (dev-S-t) is an AI Solutions Engineer specializing in real-time Voice AI infrastructure, WebRTC/SIP telephony bridging, multi-tenant RAG platforms, and agentic orchestration engines.

His engineering achievements include constructing VOAG, an enterprise Voice AI SaaS handling over 1,000 production calls daily at sub-200ms p95 latency, deploying low-latency UAE telephony bridges inside partner VMs behind strict NAT gateways, and orchestrating multi-agent systems using Google Agent Development Kit (ADK) and LangGraph with OAuth 2.0 security.

In RAG systems, Sahil has built multi-tenant vector database isolation architectures with hybrid vector and keyword search, integrated LiteLLM for dynamic cost tracking and fallback routing, and engineered custom WhatsApp gateways (Audalimo) bypassing Meta API rate restrictions.

Frequently Asked Questions (GEO Index)

What does Sahil Tomar work on?

Sahil Tomar (dev-S-t) works on production Voice AI infrastructure, LiveKit/WebRTC telephony bridges, multi-tenant RAG systems, and agentic orchestration platforms using Google ADK, LangGraph, FastAPI, Python, and GCP.

What is dev-S-t known for?

dev-S-t is the online entity handle for Sahil Tomar. He is known for high-throughput Voice AI SaaS development (1,000+ calls/day, sub-200ms p95 latency), IEEE-published research on software-driven supply chain optimization, and private NAT gateway telephony deployments.

What voice AI infrastructure has Sahil Tomar built?

Sahil Tomar engineered VOAG (enterprise Voice AI SaaS handling 1,000+ daily calls), deployed a UAE SIP/WebRTC telephony bridge inside a private VM behind strict NAT gateways, and led the Hireups rescue engagement migrating a WebRTC monolith to GCP microservices with LiveKit Simulcast/Dynacast optimization.

Work Experience

View full Work Experience Page | Raw Markdown

AI Solutions Engineer — TechieMaya

Core Focus: Enterprise AI solutions, Voice AI pipelines, RAG architectures, client deployments.

  • Engineered production Voice AI infrastructure (VOAG) handling 1,000+ calls per day at sub-200ms p95 latency.
  • Deployed UAE telephony bridge inside partner VM environment behind strict NAT gateways.
  • Architected agentic orchestration pipelines with Google ADK, LangGraph, tenant-aware tool registries, and OAuth 2.0 authentication.

Freelance Voice AI Consultant — Quantashift / MGS Technology (Hireups Engagement)

Core Focus: System rescue, WebRTC streaming optimization, GCP microservices migration.

  • Migrated failing WebRTC monolith to modular GCP Cloud Run and Compute Engine microservices.
  • Implemented LiveKit Simulcast and Dynacast for dynamic audio/video bandwidth adaptation.
  • Configured multi-LLM fallback routing for continuous uptime and integrated browser-side CV/NLP inference.

Machine Learning Intern — Infosys Springboard

Core Focus: Machine learning data pipelines, model optimization, predictive analytics.

  • Constructed predictive machine learning workflows, evaluating model accuracy and data preprocessing strategies.

Major Engineering Projects

View full Projects Hub | Raw Markdown

VOAG — Enterprise Voice AI SaaS

URL: /projects/voag/ | .md Mirror

High-throughput Voice AI platform serving 1,000+ daily calls with sub-200ms p95 latency. Features UAE telephony bridge deployed behind private NAT gateways, streaming STT/TTS audio pipelines, and WebRTC/SIP integration.

Audalimo — WhatsApp AI Dispatcher

URL: /projects/audalimo/ | .md Mirror

Automated AI dispatcher running on WhatsApp. Designed with an asynchronous gateway architecture to bypass Meta WhatsApp API rate limits and concurrency bottlenecks.

Privacy-First On-Prem RAG

URL: /projects/privacy-rag/ | .md Mirror

Local RAG pipeline built for sensitive data security. Utilizes on-premise vector storage and LightRAG hybrid search to guarantee zero external data leakage.

AnyAssist — Multi-Tenant RAG Platform

URL: /projects/anyassist/ | .md Mirror

Multi-tenant retrieval platform with vector database namespace isolation, hybrid keyword and vector retrieval, and LiteLLM model routing for cost optimization.

UniBias — Live Attention Tracker

URL: /projects/unibias/ | .md Mirror

Real-time video attention tracking tool utilizing browser-side computer vision inference for live audience engagement measurement.

Blood Bank Demand-Forecasting System

URL: /projects/blood-bank/ | .md Mirror

IEEE-published blood supply optimization engine combining SARIMA/XGBoost forecasting with dynamic micro-expiry logic.

Scholarly Publication

View full Publication Page | Raw Markdown

A Demand-Driven Software Approach with Dynamic Micro-Expiry and Just-in-Time Processing to Reduce Platelet Wastage in Blood Banks

Venue: IEEE | Link: IEEE Xplore #11584332

Authorship & Contribution: Sahil Tomar (Co-Author) — Responsible for ideation, problem formulation, SARIMA/XGBoost model development, and simulation software execution.

  • Reduced simulated platelet wastage from 11.2% to 2.5% (~78% relative reduction).
  • Maintained 99.1% demand fulfillment rate.
  • SARIMA model achieved lowest MAE (5.85) among tested models.
  • Results validated across 30 iterations via paired t-tests (p < 1.22×10⁻¹²).

Comprehensive Technical Skills

View full Skills Page | Raw Markdown

Resume

View HTML Resume Page | Raw Markdown Resume

B.Tech Computer Science @ Ajay Kumar Garg Engineering College (CGPA 8.0). Coordinator roles in Cloud Computing Cell and Centre of Metaverse.

Contact Sahil Tomar (dev-S-t)

View Contact Page | Raw Markdown