Deepak Kumar Jha

AI Demos — Senior Full-Stack Engineer · Technical Lead

Three systems, live on my own infrastructure, each with a one-click demo account against its real API — no signup, nothing mocked. This is the difference between claiming production AI experience and handing you a URL where it's running.

Privacy-First RAG

Live

Retrieval-Augmented Generation on self-hosted Llama 3 and Mistral via Ollama. Built for cases where data cannot leave the client’s infrastructure — regulated finance, government, healthcare. Upload a PDF, ask questions, get answers with citations and per-passage similarity scores.

Node.js NestJS Angular Ollama Llama 3 Qdrant (vector search)

Multi-Model Routing

Live

Routes each request between cloud LLMs (Gemini, OpenAI) and local inference based on cost, latency and data sensitivity — and shows the decision, the fallback chain and the money saved, per request and on a live dashboard.

Node.js NestJS Angular PostgreSQL Ollama Cloud LLM APIs

Agentic Workflows (MCP)

Live

A Model Context Protocol agent connected to a read-only SQL database, a live weather API and a calculator — answering multi-step questions with every tool call, result and recovery from failure streamed to the screen as it happens.

Model Context Protocol Node.js NestJS Angular PostgreSQL Ollama

A production LLM feature already shipped in a live product — see the

SmartSitting case study →
Available for new opportunities

Have a system that needs an architect?

I'm open to full-time roles, contract engagements and product/architecture work. Everything on this site is a system I designed, built, deployed and still operate — the next one could be yours.

“Give me a problem. I'll architect it, build it, and ship it.”

— Deepak Kumar Jha