Get Started
All posts
Research

Simit-TR: The Road to Our Own Turkish Model

We are fine-tuning an open-weight base model on our own hardware with consented, anonymised Turkish data to build a Türkiye-specialised model.

Simit-TR is the working name of our local model, the long-term heart of Simit GPT. Our goal isn't to replace large general-purpose models, but to grow a specialist that is better, faster and cheaper than them for Turkish and Türkiye-specific work.

Approach

  1. Base model: we start from a multilingual, open-weight model family whose licence allows commercial use.
  2. Fine-tuning: efficient methods such as LoRA/QLoRA on a single RTX 3090 (24 GB).
  3. Data: only openly licensed sources, content we create, and examples users have explicitly consented to share, anonymised and human-reviewed. Details: Responsible Data.
  4. Evaluation: we are building our own evaluation set for Turkish grammar, knowledge of Türkiye, instruction following and safety. We will share results when, and only as far as, we have measured them.

How will it be served?

Simit-TR will run behind an OpenAI-compatible server (e.g. vLLM or llama.cpp) and will never be exposed to the internet. The Simit server will reach it only over a private VPN tunnel.

Status

Simit-TR is currently in the planning and data preparation phase. Simit GPT's infrastructure is designed to switch to it with a single configuration change once it is ready.