The problem
Millions of Arabic speakers, especially in Lebanon, lack quick access to reliable medical guidance in their own dialect. Existing symptom checkers are English-only, use formal medical language, and don’t understand Lebanese colloquial (“عندي وجع راس” or “3endi waja3 ras”).
Hakim takes symptom descriptions in Lebanese Arabic, Franco-Arab (3ammiye) or English, triages urgency into GREEN, YELLOW or RED, suggests possible conditions with medical citations, and recommends next steps, all behind strict safety guardrails.
- GREEN
- Self-care at home. For example a common cold or a mild headache.
- YELLOW
- See a doctor within 24 to 48 hours. For example persistent fever or ear pain.
- RED
- Emergency, go to the ER now. For example chest pain with arm numbness, or difficulty breathing.
How a message flows
Every message goes through the same path. The safety gate sits before retrieval and the model, so a blocked request never reaches either.
Understand the dialect
An Arabic processor extracts symptoms from a 246-term Lebanese medical lexicon, so Lebanese Arabic and Franco-Arab input both map to known symptoms.
Safety gate
Pre-model filters decide whether the request is allowed at all. If it is blocked, the user gets a refusal with a referral instead of an answer.
Retrieve medical knowledge
A RAG pipeline does multi-query retrieval with reranking over a ChromaDB knowledge base, so the answer can cite its sources.
Triage
A five-step triage engine classifies GREEN, YELLOW or RED, including compound emergency patterns such as chest pain together with arm numbness.
Answer in Lebanese Arabic
The LLM writes the reply in Lebanese Arabic. Groq is the primary provider and Gemini the backup; which goes first is a setting (
LLM_PRIMARY).Post-process and stream
The reply is sanitised, a disclaimer is injected, and the result streams to the browser over server-sent events.
Safety rules
Hakim is not a diagnostic tool; it gives triage guidance only. These are the hard rules, and they are never bypassed:
- Never gives a diagnosis. It always says “possible conditions.”
- Never recommends specific medications or dosages.
- Detects emergencies from compound symptom patterns, such as chest pain together with arm numbness.
- Refuses to triage infants under 2 years, pregnancy complications, and lab results.
- Injects a disclaimer on every response.
Evaluation
A 50-scenario suite covers GREEN, YELLOW, RED and edge cases, and it runs in CI. The targets are the ones the README sets; they are goals the pipeline is gated on, not a claim about results.
- Emergency recall
- Target of at least 95%. A RED emergency must be caught.
- False alarm rate
- Target of at most 10%. GREEN cases wrongly escalated to RED.
- Triage accuracy
- Tracked: the overall share of correct classifications.
- Response quality
- Scored 1 to 5 by an LLM judge: relevance, empathy and safety.
The CI pipeline enforces the 95% emergency-recall gate. Tracing and cost tracking run through Langfuse, so latency, cost per call and error rate can be watched in production.
Stack
- Backend
- FastAPI on Python 3.11: async-first, with native SSE streaming
- Frontend
- React 19, TypeScript and Tailwind CSS v4, RTL-ready, Arabic and English
- LLM
- Groq (primary) and Gemini (backup) behind one client with caching and retry
- Embeddings
- Gemini
text-embedding-004, 768 dimensions - Vector store
- ChromaDB, local and zero-config, with hybrid search
- Observability
- Langfuse for LLM tracing and cost tracking
- Hosting
- Vercel (frontend) and Render (backend), auto-deployed from GitHub
- CI/CD
- GitHub Actions: lint, test, the eval pipeline, and deploy on merge
- FastAPI
- ChromaDB
- Gemini API
- Groq
- React
- Langfuse
