
The focus of this project was to understand how a retrieval-augmented generation (RAG) system works, from document ingestion through cloud deployment.
- Built an offline ingestion pipeline with PyMuPDF, recursive chunking, SentenceTransformer embeddings, a FAISS vector index, and SQLite source metadata.
- Developed a FastAPI query service that retrieves relevant chunks, constructs grounded prompts, validates structured Hugging Face responses, and maps citations back to application-controlled document metadata.
- Added query and sensitive-output guardrails, insufficient-context handling, request tracing, readiness checks, and deterministic unit, integration, API, and end-to-end tests.
- Evaluated retrieval with Hit Rate@K, Recall@K, Mean Reciprocal Rank, and retrieval latency using a versioned question dataset.
- Created a Streamlit client and reproducible local environment with Docker Compose, then deployed the two-container application to Google Cloud Run.
- Automated verification and manual releases with GitHub Actions, Workload Identity Federation, Artifact Registry, Secret Manager, and OpenTofu with remote state in Google Cloud Storage.
The project has two different stages. First, it constructs an offline index. Then, the online pipeline retrieves relevant passages and answers user questions in real time.
PDF corpus
-> extract and chunk text
-> generate embeddings
-> persist FAISS index and SQLite metadata
User question
-> guard and embed question
-> retrieve relevant chunks
-> build grounded prompt
-> generate and validate answer
-> return application-verified citations
I used a layered architecture to keep the domain models and application workflows independent from concrete tools such as FAISS, SQLite, and Hugging Face. This made it possible to test the core behavior with deterministic fakes while keeping hosted LLM calls and production credentials outside the test suite.
For deployment, I chose Google Cloud to apply concepts I already knew from AWS in a different cloud ecosystem (and it gave $300 free credit to use :D). The current setup runs Streamlit and FastAPI as two containers in one private Cloud Run service, stores images in Artifact Registry, loads the LLM token from Secret Manager, and uses GitHub OIDC credentials. I also incorporated a CI/CD pipeline into the repo.
DEPLOYMENT
GitHub Repository
|
v
GitHub Actions
| |
| +---- OIDC authentication ----> Google Cloud
|
+---- Build container image ----> Artifact Registry
|
v
+---------------------+
User Browser ---- IAP -------> | Cloud Run (private) |
| |
| Streamlit :8080 |
| | |
| v |
| FastAPI :8000 |
+----------+----------+
|
v
Hugging Face API
Secret Manager ---- HuggingFace token ---> FastAPI
OpenTofu + GCS state --------------------> Cloud Run configuration
Future improvements would include adding a reranking structure to improve retrieval, making the infrastructure code more complete, and deploying separate Cloud Run services so the API backend can stay private. But for now, I think I am happy with where I am…