Having a unified API across LLM was definitely a big win. We can switch between GPT-4, Claude and internal models without touching the code.Semantic Caching and fallback routing kept latency low and avoided issues when latency slowed down.
April 22, 2026
Could be more proactive in highlighting downtimes at their end
February 11, 2025