Building homegrown AI means working with local data, and that is a big responsibility. Before we start training Simit-TR, we are writing our principles down.
Our principles
- No by default. Your conversations are not used for model training unless you explicitly allow it.
- Separate, revocable consent. Training consent isn't buried in the terms of use; it's a separate choice you can withdraw at any time.
- Anonymisation. Every example flagged as a training candidate is scrubbed of personal data such as names, national ID numbers, phone numbers, emails and addresses.
- Human review. No example enters a training set without being reviewed for quality and suitability.
- Traceability. For each example we record which model produced the answer, user feedback, any corrected answer and the review status.
- KVKK compliance. We process personal data in line with Türkiye's Personal Data Protection Law No. 6698 (KVKK). Details are in our Privacy Policy.
In practice
In Simit GPT's data model, every training candidate carries consent status, anonymisation status and quality review fields. An example only counts as eligible when all three are satisfied, and that rule is defined in code.
Today
Only question-answer pairs you rate with 👍/👎 are kept as training candidates, with personal details masked, and they go through human review.