AI Software Engineer · LLM · RAG · Agentic Systems

Bhragu Gour

 

~1.8 years shipping and debugging production backend services and REST APIs, focused on LLM-driven systems. I built an end-to-end, evaluation-gated LLM platform in Python / FastAPI — multi-turn agentic tool-calling, RAG, Redis caching and session state, PostgreSQL (pgvector), and output guardrails with citation-grounding and abstention that fail closed to prevent unverified disclosure. Strong on async, event-driven patterns, schema design and query performance, and token cost / latency optimization.

bhragu@docuchunk: ~
$ whoami --verbose

Backend engineer who moved into LLM systems

Applied AI Engineer specializing in LLM and RAG systems, on production-debugging foundations.

I'm a Software Engineer — Backend & AI Integrations at Think Exam (a Ginger Webs company), a multi-tenant SaaS online examination platform. The work is latency- and correctness-critical and compliance-sensitive: sessions are live, high-stakes, and there is no second attempt for a candidate whose exam broke.

My AI surface there is proctoring. I integrated third-party AI detection pipelines — face, gaze and anomaly detection — into the production platform over REST APIs and webhooks, building the event ingestion, async processing and violation-flag persistence behind them. Alongside that I design and optimize REST APIs in Laravel, where query optimization and indexing over large datasets improved response times by 30%+, migrated a legacy PHP codebase to Laravel 10 with middleware, RBAC and Sanctum token auth, and owned the operational side — production debugging on Linux, log analysis, precise minimal-footprint fixes for high-priority incidents. I was recognized as Engineer of the Year for consistent delivery and end-to-end ownership.

Integrating someone else's model taught me the shape of these systems from the outside. So I built the inside. DocuChunk is an end-to-end, evaluation-gated LLM / RAG and agentic platform I designed and wrote myself — hybrid retrieval (dense + BM25 + RRF) with cross-encoder reranking, a multi-turn agentic ReAct tool-calling loop with five guardrails, three interchangeable vector backends, and an LLM verification layer with citation-grounding and abstention.

The part I care most about is the one most projects skip: the quality layer. A RAGAS-aligned evaluation harness (faithfulness, answer-relevancy, context precision/recall) sits behind a hermetic CI regression gate, with Langfuse tracing across 21 span types — so a prompt or model change is measured before release rather than discovered in production.

nameBhragu Gour
titleAI Software Engineer
currentSWE — Backend & AI Integrations
companyThink Exam · Ginger Webs
experience~1.8 yrs production backend
focusLLM · RAG · agentic tool-calling
corePython · FastAPI · PostgreSQL · Redis
educationB.Tech Big Data Analytics, CU
award★ Engineer of the Year
status● open to opportunities
$ cat eval/results.json

Numbers, not adjectives

Every figure below came out of a harness in the repository, on a fixed benchmark.

MRR · retrieval
0.78→0.97
Hybrid retrieval (dense + BM25 + RRF) with cross-encoder reranking, on a 42-document distractor benchmark.
Faithfulness · generation
0.71→0.89
LLM verification layer — citation parsing/validation, groundedness checks and abstention, designed to fail closed.
Context-recall · agent
+0.43
Agentic ReAct retrieval loop (tool-use + 5 guardrails) over vanilla RAG under scarce retrieval budgets.
Cost · per query
~₹0.09
Per-query cost attribution, sliced by stage / model / backend — the input to token cost and latency tuning.
16K
LOC · DocuChunk
510
Passing tests
21
Langfuse span types
30%+
API response-time gain
$ cat ./what-i-do.md

What I actually work on

Six areas I've shipped in — not a checklist of things I've read about.

Generative AI / LLM
LLM application development, prompt engineering for multi-turn interactions, agentic AI (ReAct, tool-calling) with guardrails, semantic search and embeddings, SSE streaming, semantic caching, token / cost optimization.
RAG · ReAct · guardrails
Retrieval & vector DBs
pgvector, ChromaDB and Pinecone behind one interface with per-document routing and cross-backend query merge. Hybrid search (dense + BM25 + RRF), cross-encoder reranking, MMR, chunking strategies.
pgvector · Chroma · Pinecone
Evaluation & observability
RAGAS-aligned metrics — faithfulness, answer-relevancy, context precision/recall — plus LLM-as-judge, Langfuse tracing and CI regression gates, so prompt and model changes are measured before release.
RAGAS · Langfuse · CI gates
Async & event-driven serving
Webhooks, queues, background / async workers (arq), SSE token streaming and event ingestion pipelines — including third-party AI detection events landing on a live, high-stakes platform.
webhooks · arq · SSE
Backend, data & auth
Python and PHP (Laravel), REST API design, PostgreSQL schema design with indexing and query optimization, MySQL, Redis caching and session state, RBAC, JWT / OAuth2 and session auth on multi-tenant SaaS.
FastAPI · PostgreSQL · Redis
DevOps & foundations
Docker and Docker Compose, GCP (compute, storage, managed Postgres), GitHub Actions CI/CD, Git, Linux, Postman — on Data Structures, OOP, DBMS and design-pattern foundations.
Docker · GCP · GitHub Actions
$ git log --author="Bhragu Gour" --reverse

Experience & education

Jul 2025 — Present
Software Engineer — Backend & AI Integrations
Think Exam — A Ginger Webs Company · multi-tenant SaaS
  • Integrated third-party AI detection pipelines (face / gaze / anomaly proctoring) into a production multi-tenant platform via REST APIs and webhooks — event ingestion, async processing and violation-flag persistence for live, high-stakes, latency- and correctness-critical sessions.
  • Designed and optimized REST APIs with relational-DB query optimization and indexing (joins, aggregations over large datasets), improving response times by 30%+.
  • Owned the operational side of shipped features — production debugging on Linux, log analysis, issue reproduction and precise, minimal-footprint fixes for high-priority incidents.
  • Migrated a legacy PHP codebase to Laravel 10; implemented middleware, RBAC and token-based auth (Laravel Sanctum) for compliance-sensitive access, and automated workflows with cron / background workers.
  • Recognized as Engineer of the Year for consistent delivery and end-to-end ownership across test-engine, reporting, dashboard and proctoring modules.
Jan 2025 — Jun 2025
Software Engineer Intern
Think Exam — A Ginger Webs Company
  • Built backend modules and APIs (PHP, Laravel, MySQL) for candidate management, test workflows and reporting.
  • Improved performance via query optimization and caching, and helped resolve production issues.
2021 — 2025
B.Tech — Big Data Analytics
Chandigarh University · CGPA 7.73 / 10
Data Structures, OOP, DBMS and design patterns — the foundations the rest of this page is built on.
$ pip list --tools-i-ship-with

Technical skills

Languages & backend
PythonPHP (Laravel)FastAPI REST API designasyncioOOP Design PatternsDSA
LLM & agentic AI
RAGPrompt engineeringTool-calling ReAct agentsGuardrailsMulti-turn state Semantic searchEmbeddingsSemantic caching SSE streamingToken / cost optimization
LLM providers & frameworks
Google GeminiOpenAI-compatibleOllama Hugging FaceLangChain (LCEL)LlamaIndex sentence-transformerscross-encoders
Retrieval & databases
PostgreSQLpgvectorChromaDB PineconeMySQLRedis Hybrid search (dense + BM25 + RRF)Cross-encoder reranking MMRChunking strategiesIndexing & query optimization
Evaluation & observability
RAGAS-aligned metricsFaithfulnessAnswer-relevancy Context precision / recallLLM-as-judge Langfuse tracingCI regression gates
Async, cloud & DevOps
WebhooksQueuesarq workers Event ingestionDockerDocker Compose GCPGitHub Actions (CI/CD)Git LinuxPostman
Security & compliance
RBACJWT / OAuth2Session auth Laravel SanctumMulti-tenant SaaS HIPAA-aware design
$ ./docuchunk --pitch

Featured project — the one you can actually open

Everything above is a claim. This is the running system behind it.

Featured project · DocuChunk

Talk to your documents.
Get answers you can verify.

An end-to-end RAG platform in Python / FastAPI — 16K LOC, 510 passing tests. Upload a PDF or DOCX, ask in plain language, and get an answer with inline [n] citations — produced by hybrid retrieval, cross-encoder reranking and a multi-turn agentic tool-calling loop, then checked by a verification layer that fails closed rather than let an unverified answer through.

Hybrid retrieval that measurably works
Dense vectors + BM25 + Reciprocal Rank Fusion + cross-encoder reranking — MRR 0.78 → 0.97 on a 42-document distractor benchmark.
An agent with brakes
Multi-turn ReAct tool-calling with 5 guardrails — step cap, loop detection, fallback-to-RAG — worth +0.43 context-recall under scarce retrieval budgets.
Output guardrails that fail closed
Citation parsing/validation, groundedness checks and abstention — faithfulness 0.71 → 0.89, so a verifier outage never surfaces as a confident, unverified answer.
Pluggable everywhere it matters
Three vector backends (pgvector / ChromaDB / Pinecone) with per-document routing and cross-backend merge; four LLM providers behind one protocol.
Evaluation-gated by design
RAGAS-aligned harness + hermetic CI regression gate + Langfuse tracing across 21 span types, with per-query cost attribution at ~₹0.09.
live query path
01Classify querytype · scope
▼
02Dense + BM25RRF fusion
▼
03Cross-encoder rerankprecision
▼
04Agentic ReAct loop5 guardrails
▼
05Grounded generation[n] cites
▼
06Verify or abstainfails closed
Want the whole thing, not the summary?

The live platform is one click away — the full feature set, the pipeline walkthrough, the privacy model, the stack, and a free account you can upload your own document to.

Go to DocuChunk
$ mail -s "hello" bhragu

Let's talk

Happy to walk through any layer of this — retrieval, the agent loop, the guardrails, or how the CI regression gate decides a change isn't good enough to ship.