How good is the categoriser?
The categoriser is run over a labelled synthetic set. Ground truth comes from the Demo Bank generator and a hand-authored hard-case set — independent of the model, so the score is not circular. Results below are the deterministic baseline; run the live Gemini path to compare.
Categoriser accuracy
Rules baseline1059 / 1067 labelled transactions correct across 16 categories.
Each click runs the live model (Groq GPT-OSS 120B, with Gemini 2.5 Flash failover) over a 40-line sample. The free-tier providers are rate-limited, so if a fresh call is not available this falls back to the model's cached labels from an earlier run(still the model's own output), and only to the deterministic rules baseline if nothing is cached. The model only labels transactions — the credit decision itself is deterministic. Synthetic data throughout.
Per-category precision & recall
| Category | Support | Precision | Recall | F1 |
|---|---|---|---|---|
| Salary | 42 | 100% | 98% | 99% |
| Rent | 45 | 100% | 98% | 99% |
| Utilities | 119 | 100% | 100% | 100% |
| Insurance | 37 | 100% | 97% | 99% |
| Groceries | 412 | 100% | 100% | 100% |
| Transport | 102 | 100% | 100% | 100% |
| Subscriptions | 44 | 96% | 100% | 98% |
| Loan repayment | 25 | 100% | 96% | 98% |
| BNPL | 20 | 100% | 95% | 97% |
| Other credit | 12 | 100% | 100% | 100% |
| Discretionary | 135 | 100% | 99% | 100% |
| Transfer | 42 | 95% | 98% | 96% |
| Fees | 19 | 100% | 100% | 100% |
| Gambling | 13 | 100% | 92% | 96% |
Confusion matrix
Rows = true category, columns = predicted. The diagonal is correct.
| Salary | 41 | 1 | ||||||||||||||
| Recurring income | ||||||||||||||||
| Rent | 44 | 1 | ||||||||||||||
| Utilities | 119 | |||||||||||||||
| Insurance | 36 | 1 | ||||||||||||||
| Groceries | 412 | |||||||||||||||
| Transport | 102 | |||||||||||||||
| Subscriptions | 44 | |||||||||||||||
| Loan repayment | 24 | 1 | ||||||||||||||
| BNPL | 1 | 19 | ||||||||||||||
| Other credit | 12 | |||||||||||||||
| Discretionary | 1 | 134 | ||||||||||||||
| Transfer | 1 | 41 | ||||||||||||||
| Fees | 19 | |||||||||||||||
| Gambling | 12 | 1 | ||||||||||||||
| Other |
Hard cases, shown failing honestly
Hand-authored ambiguous transactions where categorisation is genuinely hard. 8 of 14 are misclassified by the current path — we show them rather than hide them.
Irregular salary paid as a personal transfer — no 'Gehalt' keyword to anchor on.
A BNPL instalment that calls itself an 'Abo' — reads like a subscription.
An informal loan repaid to a family member — looks like an ordinary transfer.
Gambling top-up behind a generic 'entertainment' merchant name.
Legal-protection insurance billed as a 'membership' — overlaps with subscriptions.
Rent to a private landlord via standing order, without the word 'Miete'.
A café inside a supermarket brand — eating out, not groceries.
A returned rental deposit — a credit, but NOT income.
The limitation we own
Ground-truth labels here come from the Demo Bank generator and a hand-authored hard-case set — they are independent of the categoriser, which keeps the score from being circular, but they measure consistency against a known synthetic standard, not absolute truth. A real evaluation needs human-reviewed labels with measured inter-rater agreement, a representative sample of live data, and per-category cost weighting (mislabelling a loan repayment matters more than a café). The number on this page is a development signal, not a production claim.