~1.8 years shipping and debugging production backend services and REST APIs, focused on LLM-driven systems. I built an end-to-end, evaluation-gated LLM platform in Python / FastAPI — multi-turn agentic tool-calling, RAG, Redis caching and session state, PostgreSQL (pgvector), and output guardrails with citation-grounding and abstention that fail closed to prevent unverified disclosure. Strong on async, event-driven patterns, schema design and query performance, and token cost / latency optimization.
Applied AI Engineer specializing in LLM and RAG systems, on production-debugging foundations.
I'm a Software Engineer — Backend & AI Integrations at Think Exam (a Ginger Webs company), a multi-tenant SaaS online examination platform. The work is latency- and correctness-critical and compliance-sensitive: sessions are live, high-stakes, and there is no second attempt for a candidate whose exam broke.
My AI surface there is proctoring. I integrated third-party AI detection pipelines — face, gaze and anomaly detection — into the production platform over REST APIs and webhooks, building the event ingestion, async processing and violation-flag persistence behind them. Alongside that I design and optimize REST APIs in Laravel, where query optimization and indexing over large datasets improved response times by 30%+, migrated a legacy PHP codebase to Laravel 10 with middleware, RBAC and Sanctum token auth, and owned the operational side — production debugging on Linux, log analysis, precise minimal-footprint fixes for high-priority incidents. I was recognized as Engineer of the Year for consistent delivery and end-to-end ownership.
Integrating someone else's model taught me the shape of these systems from the outside. So I built the inside. DocuChunk is an end-to-end, evaluation-gated LLM / RAG and agentic platform I designed and wrote myself — hybrid retrieval (dense + BM25 + RRF) with cross-encoder reranking, a multi-turn agentic ReAct tool-calling loop with five guardrails, three interchangeable vector backends, and an LLM verification layer with citation-grounding and abstention.
The part I care most about is the one most projects skip: the quality layer. A RAGAS-aligned evaluation harness (faithfulness, answer-relevancy, context precision/recall) sits behind a hermetic CI regression gate, with Langfuse tracing across 21 span types — so a prompt or model change is measured before release rather than discovered in production.
Every figure below came out of a harness in the repository, on a fixed benchmark.
Six areas I've shipped in — not a checklist of things I've read about.
Everything above is a claim. This is the running system behind it.
An end-to-end RAG platform in Python / FastAPI — 16K LOC, 510 passing tests. Upload a PDF or DOCX, ask in plain language, and get an answer with inline [n] citations — produced by hybrid retrieval, cross-encoder reranking and a multi-turn agentic tool-calling loop, then checked by a verification layer that fails closed rather than let an unverified answer through.
The live platform is one click away — the full feature set, the pipeline walkthrough, the privacy model, the stack, and a free account you can upload your own document to.
Go to DocuChunkHappy to walk through any layer of this — retrieval, the agent loop, the guardrails, or how the CI regression gate decides a change isn't good enough to ship.