SYSTEM ARCHITECTURE · v3.0

kanoon.live

Web · Mobile · WhatsApp — Indian Legal Intelligence Platform

6
SERVICES
861
ACTS
35k+
SECTIONS
3
CLIENTS
01 · SYSTEM LAYERS
CLIENT LAYER
◻Web App
Next.js 14 · Vercel · SSR for SEO
◻Mobile App
React Native · Expo · Push Alerts
◻WhatsApp
Meta Cloud API · Webhook · OTP Login
API GATEWAY
◈API Gateway
Nginx · Rate Limiting · Auth Middleware
◈Auth Service
Supabase Auth · JWT · OTP for WhatsApp
CORE SERVICES
◆Semantic Search
FastAPI · BGE-M3 + Vyakyarth dual-encoder · bge-reranker-v2-m3 · RRF fusion
◆Document Analysis
FastAPI · PDF Parser · Section-aware Chunker
◆LLM Reasoning
Claude Sonnet (default) · DeepSeek V4-Flash (cost alt) · RAG Pipeline · Citation Engine
◆Case Tracking
FastAPI · eCourts Scraper · Alert Dispatcher
◆Acts Library
861 Acts · 35k+ Sections · REST API
◆WhatsApp Bot
Webhook Handler · Intent Parser · Session Manager
ML / RETRIEVAL LAYER
◉BGE-M3 (Corpus + Query Encoder)
BAAI · 568M params · 8192-token context · 1024-dim · MIT · Natively multilingual — encodes acts, sections and queries in Hindi/Marathi/Tamil/Telugu/Kannada/Bengali alike
◉Hybrid Retrieval + RRF Fusion
BGE-M3 Dense + Sparse (SPLADE) · Reciprocal Rank Fusion
◉bge-reranker-v2-m3
Cross-encoder · Top-50 → Top-5 · MIT · enabled as capacity allows
◉Lang Detection + Routing
langdetect · Devanagari ↔ Roman normalisation · routes queries within BGE-M3's multilingual space
ASYNC LAYER
○Job Queue
Celery · Upstash Redis broker
○Workers
Doc Embedding · Case Status Polling · Alert Dispatch
○Scheduler
Celery Beat · Daily eCourts polling for all tracked cases
DATA LAYER
▣PostgreSQL
Supabase · Users · Cases · Alert Prefs · Row-level security
▣Qdrant Cloud
Hybrid index · Dense + Sparse vectors · Acts + User Doc namespaces
▣File Storage
Supabase Storage · User-uploaded PDFs · Signed URLs
▣Redis Cache
Upstash · Session state · Search result cache · Job broker
EXTERNAL APIs
◇Anthropic Claude API
claude-sonnet-4 · Long-context RAG · Citation-aware generation
◇eCourts / NIC API
Case status · CNR lookup · Hearing dates
◇Meta WhatsApp Cloud API
Inbound webhook · Outbound templates · Media messages
↑ click any layer to highlight
02 · DATA FLOWS
select a flow above to trace the request path
03 · TECH STACK DECISIONS
Why BGE-M3 for corpus encoding (not Vyakyarth)?
BGE-M3 handles 8,192 tokens — a single Indian statute section can exceed 512 tokens, which is Vyakyarth's hard limit. BGE-M3 also emits dense + sparse + ColBERT vectors from one model for hybrid retrieval. It encodes your acts and sections at index time. Vyakyarth never touches the corpus.
Why Vyakyarth for Indic query encoding?
Vyakyarth is purpose-built for Indic languages by Ola Krutrim, built on XLM-R with contrastive fine-tuning. On the IndicXTREME Flores benchmark: Hindi 99.9, Marathi 98.8, Kannada 99.2, Malayalam 98.7, Tamil 97.9, Telugu 97.5 — decisively beating MuRIL, IndicBERT, and jina-v3. It only encodes short user queries, so its 512-token limit is irrelevant.
Why not EmbeddingGemma?
EmbeddingGemma is a 308M on-device model for phones and tablets — 2K token context, under 200MB RAM, EdgeTPU optimised. That 2K context forces mid-section chunking on Indian statutes. It's a mobile-first model, not a server-side RAG backbone. Wrong tool for a backend pipeline.
Why bge-reranker-v2-m3 bridges the cross-lingual gap?
The reranker reads the full (Indic query, English section) pair as a cross-encoder — it handles language mismatch that bi-encoders can't. A Hindi query for 'dahej pratha' will correctly match the English text of Section 498A IPC because the reranker understands both simultaneously. This is what makes the dual-encoder architecture actually work.
Why Claude Sonnet over DeepSeek V4 at launch?
Claude leads on hallucination benchmarks — in legal output, a fabricated section number is a trust-destroying bug. DeepSeek V4 routes through Chinese infrastructure, which is a data governance concern for user-uploaded case documents. DeepSeek V4-Flash is a cost alternative worth revisiting at scale once the data tradeoff is acceptable.
Why hybrid retrieval (dense + sparse)?
Indian legal text is litigated on exact terms: Section 498A, habeas corpus, vakalatnama. Pure dense retrieval misses lexical precision on proper nouns and section numbers. BGE-M3's SPLADE sparse head gives BM25-like recall without a separate index. Hybrid beats dense-only by 2–4 nDCG@10 on MIRACL.
Why Hetzner CX31, not CX21?
BGE-M3 + Vyakyarth + bge-reranker-v2-m3 together need ~7GB RAM minimum. CX21 is 4GB — won't hold all three in memory. CX31 gives 8GB for €9/mo, still 5–10× cheaper than equivalent AWS. No GPU needed at low traffic.
Why Qdrant over Pinecone?
Qdrant natively supports BGE-M3 hybrid indexing and RRF fusion across multiple retrieval channels. Pinecone doesn't. Open source, self-hostable when you outgrow the cloud tier.
Why Next.js for Web?
SSR is non-negotiable for legal search — acts and sections need to be indexable by Google. Static generation for the library pages, server rendering for search results.
Why Celery for async jobs?
PDF parsing + BGE-M3 embedding for a 100-page judgment takes 20–40 seconds. Celery queues it, the user sees a processing state, gets notified when done. Workers scale independently from the API.
Why Supabase over raw Postgres?
Auth, row-level security, and file storage are all bundled. Row-level security isolates user documents at the DB level, not just application level. Saves building 3 separate services at launch.
04 · LAUNCH COST ESTIMATE
SERVICE
COST
NOTE
Hetzner CX31 (Backend)
€9/mo
8GB RAM — FastAPI + BGE-M3 + Vyakyarth + bge-reranker + Celery workers
Vercel (Frontend)
Free
Next.js 14 · generous free tier
Supabase
Free → $25
Postgres + Auth + File Storage + Row-level security
Qdrant Cloud
Free → $25
Hybrid dense+sparse index · 1GB free tier
Upstash Redis
Free
Celery broker + session cache · serverless
Anthropic Claude API
Pay per use
~$0 at launch · scales with queries
Meta WhatsApp Cloud API
Free tier
1000 service conversations/month free
Expo (Mobile builds)
Free
React Native · OTA updates · push notification infra
TOTAL AT LAUNCH
~€9–15 / month
05 · SCALING MILESTONES
0 → 500 users~€9/mo
Hetzner CX31 · BGE-M3 + Vyakyarth + Reranker on CPU · Supabase free · Qdrant free tier · Claude Sonnet API
trigger: Launch
500 → 5k users~€80/mo
Hetzner CX41 (16GB RAM) · Supabase Pro · Qdrant paid · BGE-M3 quantized to FP16 · Evaluate DeepSeek V4-Flash as Claude cost fallback
trigger: Latency + DB storage limits
5k → 50k users~€350/mo
Hetzner dedicated + GPU node · Multiple Celery workers · Redis cluster · CDN · Benchmark Qwen3-Embedding-4B vs fine-tuned BGE-M3
trigger: Concurrency ceiling + retrieval quality plateau
50k+ usersVariable
Kubernetes on Hetzner or AWS ap-south-1 · Self-hosted Qwen3-Embedding-8B if it outperforms BGE-M3 · Self-hosted LLM when token costs justify GPU infra
trigger: Self-hosted LLM ROI + embedding model upgrade cycle
KanoonHQ · architecture plan · v3.0 · 2026BGE-M3 (corpus) · Vyakyarth (Indic queries) · bge-reranker-v2-m3 · Claude Sonnet · Qdrant hybrid