General AI models are large, costly and premium. Prism is Bitstric's research program on specialised Open Weight models. Structure is enforced at serving time; correctness is what we train for and measure. We publish results only after they are measured on a held-out benchmark.
Six steps, in order. Shown here as a fixed script.
What we are working to prove, and how we will measure it.
We train on primary regulatory text first and use synthetic examples only where needed, generated by an openly licensed model we host ourselves. Every training source is tracked back to its origin.
Output structure is enforced when the model runs, so a record either matches its schema or is not produced. Whether each field is correct is a separate question, and the one our research measures.
We are designing for compact cloud instances and local hardware. Every smaller build must re-pass the full evaluation before it is released.
A cited regulatory clause must exist. We measure how often a model cites clauses that do not, and treat it as a failure.
Real compliance inputs are often incomplete. We train for an explicit “insufficient evidence” answer and measure whether it is used correctly.
Pre-authored example output. Not generated live.
Fictitious incident report: an API endpoint exposed the records of about 500 customers for three days. The team became aware on 12 March. Containment is complete. The report does not say which categories of personal data were included.
{
"record_type": "breach_notification",
"notify_supervisory_authority": true,
"notification_window_hours_from_awareness": 72,
"citations": [
"GDPR Art. 33(1)",
"GDPR Art. 33(3)"
],
"data_subject_notification": "assess against GDPR Art. 34",
"fields_flagged_insufficient_evidence": [
"categories_of_personal_data"
]
}Design, corpus build and evaluation-set build. No trained model.
A small-scale proxy is trained and evaluated on local hardware before any cloud training spend. No bypass.
A trained candidate is measured against every evaluation measure on a held-out set. Results are recorded internally.
Named research collaborators get gated access under a research-use licence. At least one smaller build re-passes full evaluation.
Target: Q1 2027
First public release, with a model card, dataset datasheet, schema pack, a chosen licence and a cleared name.
Every evaluation target is met and the public preview has run three months without a critical issue.
All measures run on a held-out evaluation set that never enters training. Each smaller build is measured again in full. Values are published only after they are measured.
Share of outputs that parse and match their schema, with schema enforcement on at serving time.
Share of required fields whose value is correct against a reference answer.
Share of cited clauses that do not exist in the reference clause index.
How well the “insufficient evidence” answer is used on deliberately unanswerable inputs. Never answering is a failure; so is always answering.
Change in a general-knowledge benchmark versus the starting base model.
Score on a Southeast Asia compliance question set versus the prior model.
Results are reported across repeated runs, never a single run.
We are looking for research institutions and individual researchers working on these open questions. Collaboration terms are not yet fixed.
Project Prism also studies how open models can be improved safely after deployment; that work is not described here.
Research collaborators and waitlist sign-ups are welcome. We will write when there is something measured to share.