Wakeb Data Senior AI & MLOps Engineer 2024–2026

A 6-agent bilingual chatbot on self-hosted Qwen

Architected a fully async 6-agent LangGraph chatbot in Arabic and English, and ran its Qwen models on vLLM across production A100 servers.

See the architecture
Role
Architecture, build and model serving (Senior AI & MLOps Engineer) Full-time, Hybrid
Company
Wakeb Data production chatbot
When
During the Wakeb Data role Sep 2024 – Apr 2026
Where
Giza, Egypt
Stack
  • LangGraph
  • Qwen
  • vLLM
  • NVIDIA A100
  • FastAPI
  • Docker
  • Nginx
specialised agents, fully async, in LangGraph
6
GPU servers running Qwen through vLLM
A100
in one bilingual chatbot
AR + EN
Fig. 1 wakeb-multi-agent-chatbot · architecture
Multi-agent chatbot: fully async LangGraph with six specialised agentsArchitecture diagram. A user writes in Arabic or English. Requests pass through Nginx, a reverse proxy, to Dockerized FastAPI services; key internal APIs use gRPC. Inside a fully async LangGraph graph, Command-based routing with no static edges hands control between six specialised agents, built as modular subgraphs. The graph calls four back ends: the Qwen LLM served with vLLM on A100 GPUs; hybrid retrieval that combines vector and BM25 search with LLM reranking; Brave Search through MCP; and a SQLite node-level cache.LANGGRAPH · FULLY ASYNC6 AGENTSUserArabic · EnglishNginxreverse proxyFastAPI servicesDocker · gRPC1Command routingno static edges2Qwen LLMvLLM · A100 GPUs4Hybrid retrievalvector + BM25 · LLM rerank5Brave Searchvia MCP6Node cacheSQLite73orchestratoragentretrieval · datamodel · LLMservice · infrastoreexternalrequest1note
Multi-agent chatbot: fully async LangGraph with six specialised agentsArchitecture diagram. A user writes in Arabic or English. Requests pass through Nginx, a reverse proxy, to Dockerized FastAPI services; key internal APIs use gRPC. Inside a fully async LangGraph graph, Command-based routing with no static edges hands control between six specialised agents, built as modular subgraphs. The graph calls four back ends: the Qwen LLM served with vLLM on A100 GPUs; hybrid retrieval that combines vector and BM25 search with LLM reranking; Brave Search through MCP; and a SQLite node-level cache.LANGGRAPH · ASYNC6 AGENTSUserArabic · EnglishNginxreverse proxyFastAPI servicesDocker · gRPC1Command routingno static edges2Qwen LLMvLLM · A100 GPUs4Hybrid retrievalvector + BM25 · LLM rerank5Brave Searchvia MCP6Node cacheSQLite73orchestratoragentretrieval · datamodel · LLMservice · infrastoreexternalrequest1note

Figure 1 A fully async, bilingual multi-agent chatbot: LangGraph hands off between six specialised agents with Command-based edgeless routing, on Qwen served by vLLM on A100 GPUs, with hybrid retrieval, web search over MCP and node-level caching. Diagrams show the components named in public descriptions of the work. They are simplified, not complete system maps.

  1. Services. Dockerized FastAPI services behind Nginx; key internal APIs moved from REST to gRPC to cut latency between services.
  2. Routing. Command-based and edgeless: control moves between agents through LangGraph Command instead of fixed graph edges.
  3. Agents. Six specialised agents composed as modular subgraphs, working in Arabic and English.
  4. Model. Qwen served with vLLM on A100 GPU infrastructure.
  5. Retrieval. Hybrid search for RAG: vector and BM25 results, reranked by an LLM.
  6. Web search. Through the Brave Search MCP integration.
  7. Cache. SQLite-based node-level caching.
Text description of the diagram

Architecture diagram. A user writes in Arabic or English. Requests pass through Nginx, a reverse proxy, to Dockerized FastAPI services; key internal APIs use gRPC. Inside a fully async LangGraph graph, Command-based routing with no static edges hands control between six specialised agents, built as modular subgraphs. The graph calls four back ends: the Qwen LLM served with vLLM on A100 GPUs; hybrid retrieval that combines vector and BM25 search with LLM reranking; Brave Search through MCP; and a SQLite node-level cache.

On this page The problem
01

The problem

The chatbot had to answer in Arabic and English, run on self-hosted models on A100 GPU servers, ground its answers in documents and live web results, and guide users through multi-turn service configuration. In production, the model servers had to recover on their own when something failed.

02

Approach

  1. Hybrid retrieval for RAG: vector search and BM25, reranked by an LLM, with SQLite caching at the graph-node level.
  2. Live web search through the Brave Search MCP server.
  3. Served Qwen with vLLM on NVIDIA A100 GPUs and owned the lifecycle: serving configuration, GPU allocation, model versions, health monitoring and automated restarts across multiple production GPU servers.
  4. Shipped as Dockerised FastAPI services behind Nginx, with CI/CD.
03

Architecture

Component view: async LangGraph agents call self-hosted Qwen on vLLM, retrieve through hybrid search with LLM reranking, and reach the web through MCP; FastAPI services sit behind Nginx. Figure 1

Users and apps
  • Users (Arabic / English)
Orchestration
  • LangGraph orchestration · Command-based routing, async
Agents
  • 6 specialised agents · Modular subgraphs
Models
  • vLLM serving Qwen
Retrieval
  • Hybrid search: vector + BM25
  • LLM reranker
Data
  • SQLite node cache
Services
  • FastAPI services (Docker)
  • Brave Search via MCP
Infrastructure
  • Nginx
  • A100 GPU servers · Versions, health checks, auto-restart
Connections
  • Users (Arabic / English) to Nginx
  • Nginx to FastAPI services (Docker)
  • FastAPI services (Docker) to LangGraph orchestration
  • LangGraph orchestration to 6 specialised agents
  • LangGraph orchestration to SQLite node cache
  • 6 specialised agents to Hybrid search: vector + BM25
  • Hybrid search: vector + BM25 to LLM reranker
  • 6 specialised agents to Brave Search via MCP
  • 6 specialised agents to vLLM serving Qwen
  • vLLM serving Qwen to A100 GPU servers
Fig. 2 hybrid-rag · architecture
Hybrid retrieval with LLM rerankingFlow diagram. A query is sent in parallel to BM25 keyword search and to dense vector search over embeddings. Their results are merged, reranked by an LLM, and passed as context to an LLM that generates the answer.PARALLELQueryBM25keywordDense vectorembeddingsMerge2LLM rerank3ContextGenerateLLM41retrieval · datamodel · LLMexternalrequest1note
Hybrid retrieval with LLM rerankingFlow diagram. A query is sent in parallel to BM25 keyword search and to dense vector search over embeddings. Their results are merged, reranked by an LLM, and passed as context to an LLM that generates the answer.PARALLELQueryBM25keywordDense vectorembeddingsMerge2LLM rerank3ContextGenerateLLM41retrieval · datamodel · LLMexternalrequest1note

Figure 2 The retrieval pattern behind my RAG work: keyword (BM25) and dense vector search run in parallel, results are merged and reranked by an LLM, and the answer is generated from that context. Diagrams show the components named in public descriptions of the work. They are simplified, not complete system maps.

  1. Retrieve. Keyword (BM25) and dense vector retrieval run in parallel.
  2. Merge. Both result lists are merged into one candidate set.
  3. Rerank. An LLM reranks the candidates before they reach the prompt.
  4. Generate. The LLM generates the answer from the reranked context.
Text description of the diagram

Flow diagram. A query is sent in parallel to BM25 keyword search and to dense vector search over embeddings. Their results are merged, reranked by an LLM, and passed as context to an LLM that generates the answer.

Fig. 3 llm-serving-a100 · architecture
LLM serving on A100 GPUs: the model deployment lifecycleArchitecture diagram. Delivery: code in Git goes through CI/CD with shell-scripted automation into Docker images, promoted from development to production and deployed to the production GPU servers. Model version management for Qwen and the vLLM serving configuration, including GPU resource allocation, are applied to the same servers. Production: multiple A100 GPU servers, each running vLLM with a Qwen model; two are drawn and a placeholder stands for the rest. The LangGraph chatbot agents send inference requests to them. Health monitoring checks the servers, and automated restart pipelines restart them.DELIVERY · CI/CDPRODUCTION GPU SERVERSdeployGitCI/CDshell-scriptedDocker imagesdev → prodModel versionsQwen2vLLM configGPU allocation3GPU servervLLM · Qwen · A100GPU servervLLM · Qwen · A100more serversChatbot agentsLangGraphHealth monitoring5Automated restartrestart pipelines614model · LLMevaluationservice · infrastoreexternalrequestbatch · background1note
LLM serving on A100 GPUs: the model deployment lifecycleArchitecture diagram. Delivery: code in Git goes through CI/CD with shell-scripted automation into Docker images, promoted from development to production and deployed to the production GPU servers. Model version management for Qwen and the vLLM serving configuration, including GPU resource allocation, are applied to the same servers. Production: multiple A100 GPU servers, each running vLLM with a Qwen model; two are drawn and a placeholder stands for the rest. The LangGraph chatbot agents send inference requests to them. Health monitoring checks the servers, and automated restart pipelines restart them.DELIVERY · CI/CDPRODUCTION GPU SERVERSdeployGitCI/CDDockerimagesModel versionsQwen2vLLM configGPU allocation3GPU servervLLM · Qwen · A100GPU servervLLM · Qwen · A100more serversHealth monitoring5Automated restartrestart pipelines614model · LLMevaluationservice · infrastorerequestbatch · background1note

Figure 3 End-to-end deployment lifecycle for Qwen on vLLM across multiple production A100 GPU servers: CI/CD and containers, model versions, serving configuration and GPU allocation, health monitoring and automated restarts. Diagrams show the components named in public descriptions of the work. They are simplified, not complete system maps.

  1. Delivery. CI/CD and containerized deployments with Docker, Git and shell-scripted automation, across development and production.
  2. Versions. Model version management for the Qwen models in production.
  3. Serving. vLLM serving configuration and GPU resource allocation.
  4. GPU servers. Multiple production GPU servers (NVIDIA A100) running vLLM.
  5. Monitoring. Health monitoring across the GPU servers.
  6. Recovery. Automated restart pipelines.
Text description of the diagram

Architecture diagram. Delivery: code in Git goes through CI/CD with shell-scripted automation into Docker images, promoted from development to production and deployed to the production GPU servers. Model version management for Qwen and the vLLM serving configuration, including GPU resource allocation, are applied to the same servers. Production: multiple A100 GPU servers, each running vLLM with a Qwen model; two are drawn and a placeholder stands for the rest. The LangGraph chatbot agents send inference requests to them. Health monitoring checks the servers, and automated restart pipelines restart them.

04

Outcome

  • Deployed to production as a fully async, bilingual Arabic/English system.
  • Self-hosted model serving with health monitoring and automated restarts across multiple production GPU servers.
  • Answers grounded in retrieved documents and live web results.
05

Stack

Orchestration
LangGraph
Models
QwenvLLM
Retrieval
BM25Brave Search
Data
SQLite
Services
FastAPIMCP
Pipelines
CI/CD
Infrastructure
NVIDIA A100DockerNginx

Where these facts come from. Everything on this page comes from my CV, my LinkedIn profile and my GitHub profile and repositories. Nothing is estimated: where no figure is public, the page describes what the system does.

Keep reading

Questions about this build?

Ask my AI about “6-agent bilingual chatbot on vLLM”.

It answers from my CV and public profile, in English or Arabic.