Ishan Sharma is a Senior Software Engineer based in Surrey, British Columbia, Canada, specializing in production AI systems: low-latency LLM inference, retrieval-augmented generation, real-time speech pipelines, and privacy-preserving on-prem deployment. His operating principle: the model is the smallest part of the system — the deterministic execution, curated context, and trust architecture around it are what make AI dependable.
Experience
Own production agent-assist systems for contact centres: on-prem LLM summarization, real-time PII redaction, and a Ruby indexing pipeline reindexing under live traffic. Delivered sub-500ms inference at 200+ concurrent sessions on 2×H100 by tuning batching and concurrency, and built GPU instrumentation to catch degradation before it lands.
Led technical design of a decentralized container execution platform running workloads on GPU-backed DePIN networks. Defined execution flows, isolation boundaries and deployment primitives, and built orchestration to schedule and monitor across heterogeneous GPU environments.
Built citizen-facing Angular features including “Authorize My Representative,” and worked across Angular and Java services on the “File a Formal Dispute” workflow — accessibility, responsive behaviour and regression coverage for services taxpayers depend on.
Built real-time ASR pipelines on Kaldi and NVIDIA NeMo with high-availability architecture on AWS and distroless Docker images. Led the Angular 14 migration of the BI dashboard, and shipped the PII redaction system, a conversational SQL query assistant, and the desktop audio capture framework.
Server-side clickstream tracking for Shopify merchants on GCP with Tag Manager, Pub/Sub and containerized ETL — multi-tenant behavioural collection, plus async pipelines and CI/CD that cut manual setup per client.
Skills
Education & Certifications
Master of Applied Computer Science — Dalhousie University, Halifax, Canada
B.Tech, Computer Science & Engineering — Manipal University Jaipur, India
Microsoft Azure Fundamentals (AZ-900)
Canadian permanent resident
Questions
Has he handled production incidents?
Yes. When a customer reported unexplained data leaving their on-prem environment with his employer's software as the suspect, he worked logs, packages and service health in front of the customer's security team, isolated the cause to a crash-looping legacy service repeatedly pulling packages from an external host, reported it straight to the customer, and decommissioned the service. He has also restored a stalled production pipeline in about thirty minutes — a regression from his own team's update, which he says plainly — by finding a child process whose output pipes were never drained.
What is distinctive about how he builds?
He looks for the version of the problem that does not need a model: speaker attribution read from device topology instead of a diarization model; a deterministic retrieval path instead of a generated query; a single planning call compiled to a validated execution plan instead of an agent loop. Models do the work only where they earn it — everything that carries a number is deterministic code.
Who is Ishan Sharma?
Ishan Sharma is a Senior Software Engineer based in Surrey, British Columbia, Canada. He builds production AI systems — real-time LLM inference, retrieval-augmented agents, speech recognition pipelines, and on-prem PII redaction — built to be fast, private and dependable.
What does Ishan Sharma work on?
He currently owns production agent-assist systems for contact centres at Yactraq Online: on-prem LLM summarization, real-time PII redaction with 50+ configurable labels, and retrieval pipelines that hold inference under 500 milliseconds at 200+ concurrent sessions on two NVIDIA H100 GPUs.
How can I reach Ishan Sharma?
Email (ishansharma1320@gmail.com) or LinkedIn — he reads both. The conversations he enjoys most are about production AI systems that need to be fast, private and dependable.