AI inference gateway that routes prompts to the cheapest capable LLM automatically — cutting costs by 85% per request. Built with 3-tier routing (EX : Llama 8B → 70B → Qwen 32B), real-time WebSocket dashboard, Battle Mode and self-healing fallback system.