The problem
The chatbot had to answer in Arabic and English, run on self-hosted models on A100 GPU servers, ground its answers in documents and live web results, and guide users through multi-turn service configuration. In production, the model servers had to recover on their own when something failed.
Approach
- Hybrid retrieval for RAG: vector search and BM25, reranked by an LLM, with SQLite caching at the graph-node level.
- Live web search through the Brave Search MCP server.
- Served Qwen with vLLM on NVIDIA A100 GPUs and owned the lifecycle: serving configuration, GPU allocation, model versions, health monitoring and automated restarts across multiple production GPU servers.
- Shipped as Dockerised FastAPI services behind Nginx, with CI/CD.
Architecture
Component view: async LangGraph agents call self-hosted Qwen on vLLM, retrieve through hybrid search with LLM reranking, and reach the web through MCP; FastAPI services sit behind Nginx. Figure 1
- Users and apps
-
- Users (Arabic / English)
- Orchestration
-
- LangGraph orchestration · Command-based routing, async
- Agents
-
- 6 specialised agents · Modular subgraphs
- Models
-
- vLLM serving Qwen
- Retrieval
-
- Hybrid search: vector + BM25
- LLM reranker
- Data
-
- SQLite node cache
- Services
-
- FastAPI services (Docker)
- Brave Search via MCP
- Infrastructure
-
- Nginx
- A100 GPU servers · Versions, health checks, auto-restart
- Connections
-
- Users (Arabic / English) to Nginx
- Nginx to FastAPI services (Docker)
- FastAPI services (Docker) to LangGraph orchestration
- LangGraph orchestration to 6 specialised agents
- LangGraph orchestration to SQLite node cache
- 6 specialised agents to Hybrid search: vector + BM25
- Hybrid search: vector + BM25 to LLM reranker
- 6 specialised agents to Brave Search via MCP
- 6 specialised agents to vLLM serving Qwen
- vLLM serving Qwen to A100 GPU servers
Figure 2 The retrieval pattern behind my RAG work: keyword (BM25) and dense vector search run in parallel, results are merged and reranked by an LLM, and the answer is generated from that context. Diagrams show the components named in public descriptions of the work. They are simplified, not complete system maps.
- Retrieve. Keyword (BM25) and dense vector retrieval run in parallel.
- Merge. Both result lists are merged into one candidate set.
- Rerank. An LLM reranks the candidates before they reach the prompt.
- Generate. The LLM generates the answer from the reranked context.
Text description of the diagram
Flow diagram. A query is sent in parallel to BM25 keyword search and to dense vector search over embeddings. Their results are merged, reranked by an LLM, and passed as context to an LLM that generates the answer.
Figure 3 End-to-end deployment lifecycle for Qwen on vLLM across multiple production A100 GPU servers: CI/CD and containers, model versions, serving configuration and GPU allocation, health monitoring and automated restarts. Diagrams show the components named in public descriptions of the work. They are simplified, not complete system maps.
- Delivery. CI/CD and containerized deployments with Docker, Git and shell-scripted automation, across development and production.
- Versions. Model version management for the Qwen models in production.
- Serving. vLLM serving configuration and GPU resource allocation.
- GPU servers. Multiple production GPU servers (NVIDIA A100) running vLLM.
- Monitoring. Health monitoring across the GPU servers.
- Recovery. Automated restart pipelines.
Text description of the diagram
Architecture diagram. Delivery: code in Git goes through CI/CD with shell-scripted automation into Docker images, promoted from development to production and deployed to the production GPU servers. Model version management for Qwen and the vLLM serving configuration, including GPU resource allocation, are applied to the same servers. Production: multiple A100 GPU servers, each running vLLM with a Qwen model; two are drawn and a placeholder stands for the rest. The LangGraph chatbot agents send inference requests to them. Health monitoring checks the servers, and automated restart pipelines restart them.
Outcome
- Deployed to production as a fully async, bilingual Arabic/English system.
- Self-hosted model serving with health monitoring and automated restarts across multiple production GPU servers.
- Answers grounded in retrieved documents and live web results.
Stack
- Orchestration
- LangGraph
- Models
- QwenvLLM
- Retrieval
- BM25Brave Search
- Data
- SQLite
- Services
- FastAPIMCP
- Pipelines
- CI/CD
- Infrastructure
- NVIDIA A100DockerNginx
Where these facts come from. Everything on this page comes from my CV, my LinkedIn profile and my GitHub profile and repositories. Nothing is estimated: where no figure is public, the page describes what the system does.