Field notes.
Once a month: what we are learning from live projects, which models we actually run in production, what we are deliberately leaving out - and occasionally a short methodological argument we are currently testing.
Why we have just cancelled a pilot at an OEM.
A specific project where the expectations did not match reality. Plus: three questions we put on the table before every new pilot.
Claude 4.5 in an A/B against our fine-tuned Mistral.
The result was not what we expected. We show the eval matrix and the budget question that ultimately decided the choice.
A BaFin audit: 90 minutes, no findings.
Which eight artefacts turned out to be enough - and which details we are already preparing for the next audit.
Four things that did not work for us in 2025.
A year in review without the success stories. The four things we are doing differently in 2026.
Why we do not build “AI copilots”.
The difference between a copilot and an agent is architectural, not cosmetic. We explain which part we do - and which we do not.
Tool use in production: two patterns, one anti-pattern.
Which two interaction patterns we run in the field - and why “the agent calls OpenAPI directly” breaks in practice.