CLOUD VS. EDGE AI FOR PAKISTAN’S CRIMINAL LEGAL SYSTEM: A POLICY-ORIENTED COMPARATIVE ANALYSIS OF RETRIEVAL-AUGMENTED GENERATION ARCHITECTURES
Keywords:
Retrieval-Augmented Generation, Legal AI, Cloud Computing, Edge Computing, API Rate Limiting, Round Robin Load Balancing, Hybrid Retrieval, Pakistan Penal Code, Code of Criminal Procedure, Cosine Similarity, Exact-Match Caching, BM25, Reciprocal Rank FusionAbstract
To leverage Large Language Models (LLMs) in the legal domain, three conflicting dimensions of computational efficiency, data sovereignty, and cost must be balanced. The experiments are performed on the same legal corpus of 3237 chunks (characters over 1000 per chunk, 200 chars overlap), and the same size of the vector database, making it a fair comparison of the two different architectural paradigms: (1) a local edge-computing system executed completely on CPU-constrained hardware (Intel Core i5, 6th Gen, 8GB RAM), and (2) a cloud-native distributed system utilizing commercial APIs (Google Gemini for embeddings, Groq for inference) that has been intelligently rate-limited via round-robin key rotation. Our local system has aggressive context pruning, deterministic sampling (temperature=0.0) and exact-match SQLite caching, and can deliver a warm-start latency of 105-118s on aging CPUs, with a cache-hit response of 0.1-0.5s. The cloud system provides a 3,237-chunk dataset ingestion, hybrid retrieval that combines semantic search (Cosine Similarity, threshold 0.27) with BM25 keyword matching, and is also based on the hybrid retrieval using the Reciprocal Rank Fusion method with equal weights (W_Lexical=W_Sementics=0.5); it also offers ultra-low latency (warm start 1.5–3.5s, cache hit 0.1–0.5s) for the retrieval process. We measure seven important metrics: inference latency, hallucination rate, privacy guarantees, cost per query, scalability limits, retrieval precision, and ingestion time of the dataset. The results show that local deployments guarantee absolute data privacy, but come with no recurring costs, while cloud architectures provide greater responsiveness and maintainability with the drawback of data sovereignty. This study offers a framework that is supported by empirical evidence that can guide legal AI practitioners in resource-constrained jurisdictions.














