LLM · Edge AI · Efficient Inference
Offline African-Language LLM
Built an offline LLM system designed to operate on commodity hardware while supporting African-language interaction under strict compute and memory constraints.
System flow
Metrics
- Model
- Qwen2.5-1.5B-Instruct
- RAM
- ≤ 8 GB
- Tokens/sec
- 8–12 tok/s (CPU)
- Model Size
- ~1.1 GB (Q4_K_M)
- Accuracy
- In evaluation
Research angle
How do model size, quantization, grammar constraints and hardware limitations affect accuracy and inference efficiency?