Privacy-First RAG
Ask questions against your own documents and get answers with citations back to the exact passage. Self-hosted Llama 3 via Ollama — nothing leaves the server.
- Document Processing
- Vector Search
- Private & Self-Hosted
- Source Citations
Three live AI systems — RAG, MCP Agents and Multi-Model Routing — designed, built and deployed end to end by me. Not screenshots. Not slides. Working software you can open right now.
Click, try, break them. These are real systems on a real server — not screenshots.
Ask questions against your own documents and get answers with citations back to the exact passage. Self-hosted Llama 3 via Ollama — nothing leaves the server.
An agent connected to real tools over the Model Context Protocol — read-only SQL, live weather, a calculator. Every tool call, result and recovery streams to the screen as it happens.
Routes each request between local inference and cloud models on cost, latency, quality and data sensitivity — and shows you the decision, the fallback chain and the money saved.
Different questions, same evidence. Start wherever you actually are.
Real skills, real code, working systems you can open in a tab. No guesswork about what "AI experience" means on a CV.
See the demos →CTO / FounderHow I approach a product problem end to end — the trade-offs, what I chose to build, and what I deliberately did not.
Explore architecture →Architect / Eng LeadThe decisions behind each system: retrieval strategy, tool boundaries, routing policy, failure behaviour.
View technical decisions →Hiring ManagerSomeone who owns a problem from architecture through deployment and still operates it in production afterwards.
See experience →One engineer, seven stages. The last one is the claim most portfolios can't make.
Not a technology list — what each of those technologies was actually used to finish.
Three architectures, each designed before a line was written
NestJS APIs, streaming, queueing, rate limits, migrations
Angular interfaces built for the non-technical visitor
Self-hosted Llama 3, cloud LLMs, RAG, MCP, model routing
Docker Compose, nginx, TLS, CI/CD, one VPS running all three
Designed → built → deployed → operated, by one engineer
The things recruiters and hiring managers ask before the first call.
Yes. Deepak is an immediate joiner, available within 15 days, and is open to full-time Senior Full-Stack Engineer and Technical Lead roles in Delhi NCR and Dubai/UAE. Reach out via the contact form or LinkedIn.
I'm open to full-time roles, contract engagements and product/architecture work. Everything on this site is a system I designed, built, deployed and still operate — the next one could be yours.
“Give me a problem. I'll architect it, build it, and ship it.”
— Deepak Kumar Jha