The problem
ArabsStock is an Arab library of millions of royalty-free photos, videos and vector graphics. The platform needed one assistant for very different jobs (finding assets, generating images, running database operations, moderating content) and search that finds assets by meaning and by visual similarity.
Approach
- Built the semantic search pipeline: OpenAI embeddings into a Pinecone index, filled by automated batch ingestion across millions of assets.
- Added visual-similarity search, so users can find visually similar assets across the library.
- Served image search, sound generation, instrument classification, BPM detection and audio watermarking as containerised FastAPI microservices, and added sound search for a beta sound library.
- Added agent evaluation workflows that track reliability, response accuracy and routing in production.
Architecture
Component view: a LangGraph orchestrator routes requests to specialised agents; a batch pipeline keeps the vector index current; model services run as separate containers; evaluation watches routing and accuracy. Figure 1
- Users and apps
-
- Platform users
- Orchestration
-
- LangGraph orchestrator
- Agents
-
- Semantic search
- Image generation
- Database operations
- Content moderation
- Other platform AI services
- Models
-
- OpenAI embeddings
- Retrieval
-
- Visual-similarity search
- Data
-
- Pinecone vector index
- Services
-
- Audio model services · Generation, search, instruments, BPM, watermarking
- Pipelines
-
- Batch ingestion · Millions of photos, videos, vector graphics
- Quality
-
- Agent evaluation
- Connections
-
- Platform users to LangGraph orchestrator
- LangGraph orchestrator to Semantic search
- LangGraph orchestrator to Image generation
- LangGraph orchestrator to Database operations
- LangGraph orchestrator to Content moderation
- LangGraph orchestrator to Other platform AI services
- Batch ingestion to OpenAI embeddings
- OpenAI embeddings to Pinecone vector index
- Agent evaluation to LangGraph orchestrator (routing, accuracy)
Outcome
- A production assistant whose 14 agents cover search, image generation, data operations and moderation.
- Millions of photos, videos and vector graphics indexed for semantic search, plus visual-similarity search across the library.
- A new audio line for the platform: a beta sound library with generation and search.
- Agent reliability, accuracy and routing measured in production rather than assumed.
Stack
- Orchestration
- LangGraph
- Models
- OpenAI APIvLLM
- Retrieval
- Pinecone
- Data
- PostgreSQL
- Services
- PythonFastAPI
- Infrastructure
- DockerNVIDIA A100
Where these facts come from. Everything on this page comes from my CV, my LinkedIn profile and my GitHub profile and repositories. Nothing is estimated: where no figure is public, the page describes what the system does.