Why BGE-M3 for corpus encoding (not Vyakyarth)?
BGE-M3 handles 8,192 tokens — a single Indian statute section can exceed 512 tokens, which is Vyakyarth's hard limit. BGE-M3 also emits dense + sparse + ColBERT vectors from one model for hybrid retrieval. It encodes your acts and sections at index time. Vyakyarth never touches the corpus.
Why Vyakyarth for Indic query encoding?
Vyakyarth is purpose-built for Indic languages by Ola Krutrim, built on XLM-R with contrastive fine-tuning. On the IndicXTREME Flores benchmark: Hindi 99.9, Marathi 98.8, Kannada 99.2, Malayalam 98.7, Tamil 97.9, Telugu 97.5 — decisively beating MuRIL, IndicBERT, and jina-v3. It only encodes short user queries, so its 512-token limit is irrelevant.
Why not EmbeddingGemma?
EmbeddingGemma is a 308M on-device model for phones and tablets — 2K token context, under 200MB RAM, EdgeTPU optimised. That 2K context forces mid-section chunking on Indian statutes. It's a mobile-first model, not a server-side RAG backbone. Wrong tool for a backend pipeline.
Why bge-reranker-v2-m3 bridges the cross-lingual gap?
The reranker reads the full (Indic query, English section) pair as a cross-encoder — it handles language mismatch that bi-encoders can't. A Hindi query for 'dahej pratha' will correctly match the English text of Section 498A IPC because the reranker understands both simultaneously. This is what makes the dual-encoder architecture actually work.
Why Claude Sonnet over DeepSeek V4 at launch?
Claude leads on hallucination benchmarks — in legal output, a fabricated section number is a trust-destroying bug. DeepSeek V4 routes through Chinese infrastructure, which is a data governance concern for user-uploaded case documents. DeepSeek V4-Flash is a cost alternative worth revisiting at scale once the data tradeoff is acceptable.
Why hybrid retrieval (dense + sparse)?
Indian legal text is litigated on exact terms: Section 498A, habeas corpus, vakalatnama. Pure dense retrieval misses lexical precision on proper nouns and section numbers. BGE-M3's SPLADE sparse head gives BM25-like recall without a separate index. Hybrid beats dense-only by 2–4 nDCG@10 on MIRACL.
Why Hetzner CX31, not CX21?
BGE-M3 + Vyakyarth + bge-reranker-v2-m3 together need ~7GB RAM minimum. CX21 is 4GB — won't hold all three in memory. CX31 gives 8GB for €9/mo, still 5–10× cheaper than equivalent AWS. No GPU needed at low traffic.
Why Qdrant over Pinecone?
Qdrant natively supports BGE-M3 hybrid indexing and RRF fusion across multiple retrieval channels. Pinecone doesn't. Open source, self-hostable when you outgrow the cloud tier.
Why Next.js for Web?
SSR is non-negotiable for legal search — acts and sections need to be indexable by Google. Static generation for the library pages, server rendering for search results.
Why Celery for async jobs?
PDF parsing + BGE-M3 embedding for a 100-page judgment takes 20–40 seconds. Celery queues it, the user sees a processing state, gets notified when done. Workers scale independently from the API.
Why Supabase over raw Postgres?
Auth, row-level security, and file storage are all bundled. Row-level security isolates user documents at the DB level, not just application level. Saves building 3 separate services at launch.