Get Started
All posts
Engineering

Hybrid Model Routing: Ready-made API Models and Simit-TR Together

We didn't tie Simit GPT to a single model. The Model Router sends each question to the best engine and can fail over when a provider has trouble.

Tying an AI product tightly to a single model provider is easy in the short term and risky in the long term. We designed Simit GPT to be provider-independent from day one.

Layers

text
UI → /api/chat → Knowledge retrieval → Model Router → Provider
                                             ├─ Ready-made API models
                                             └─ Simit Local (Simit-TR)
  • The UI never talks to a model provider directly; API keys only live on the server.
  • Providers implement a common interface: chat() and streamChat().
  • The Model Router runs in the mode chosen in configuration: ready-made API models, local model or hybrid.

Hybrid mode

In hybrid mode the router tries providers in order. If a provider hits a transient problem (a timeout, a network error or a rate limit) before producing its first word, the router moves on to the next one. We never switch models after an answer has started, so you never get a "spliced" response.

Later the order won't be static: Turkish-heavy and Türkiye-specific questions can go to Simit-TR, while long reasoning tasks go to strong ready-made API models.

Why it matters

  • Resilience: an outage at one provider doesn't take the product down.
  • Cost: simple tasks can be served by a cheaper model.
  • Independence: bringing in our own model is a configuration change.