How we check
An AI assigns your description to the rules of the ordinance. A fixed program recalculates hours, level and money, without AI. The result can be wrong. So in the end, you check and decide.
How it works
- You
Describe
You describe your daily life in your own words.
- AI
Assign
The AI assigns your sentences to rules, with a word-for-word quote.
- Program
Calculate
A fixed program calculates according to the ordinance.
- You
Decide
You go through every line and decide what to do with it.
How often it is right
Most reliable so far: untouched court cases
5 of 15
On court cases that we did not touch again after the build, it got the level right on the first try.
Current status: practice cases
10 of 14
On practice cases from published judgments, which we used to tune it, it got the level right on the first try.
Pflegegeld-Prüfer helps you prepare. The authority or the court sets the level. Get advice.
Preparation
We collected 29 published judgments on Pflegegeld (long-term care allowance). In 13 of these cases, the original Bescheid (official decision) was not at the level the court set in the end.
Rules fixed in advance
Before we measured for the first time, we fixed what counts as a hit and which target applies: 12 of 14 practice cases.
First measurements
The first run hit 9 of 14. We improved it once, then froze the rules and the instructions to the AI. The last run hit 10 of 14, so the target was missed.
In 4 of these hits, the right answer was no level: no entitlement, or a child for whom it does not calculate. In cases with a real level, it hit 6 of 9.
The test on untouched cases
Once, at the end, on 15 further judgments: 5 correct. After that we changed nothing. In cases with a real level, it was 4 of 13.
Since then
The site is public. We have improved the follow-up questions and the interface. On 28.09.2026 we compared 10 AI models on the same cases. Gemini 3.8 Flash, the model for your own text, hit the practice cases as often as the model we use as the benchmark for quality. On 29.09.2026 we measured again. The result is shown above as the current status.
We measure on published judgments where the right level is already known. We use one part to practise and improve. We leave another part untouched and only test it at the end.
We fix in advance what counts as a hit and write it down. Later changes are recorded there with a date. Every measurement is added and none is deleted, including the bad ones.
We keep questioning the method itself. An automatic check without AI runs with every change. It recalculates the reports, checks the denominators and looks for traces of the untouched cases in the rules. It cannot measure again. What we find is listed below under the limits.
Measured on 29.09.2026 with Gemini 3.8 Flash, the model that also reads your own text. We tuned it on these cases, so the number is probably too high.
Measured once on 26.09.2026. This number says the most about how it does on new cases.
For comparison: simply taking the level from the Bescheid in these cases would have been right in 8 of 14 cases, which is more often. So it does not reliably tell whether a Bescheid is correct.
If the program gets the hours from the judgment, the level is right in all 22 practice and made-up cases. For the untouched cases, in 12 of 15. The other judgments give no hours per task, so the program has no input.
When it was wrong, it was mostly too low: ten times too low, twice too high.
What does this mean for you?
If it shows more than your Bescheid, it is worth a close look and counselling.
If it does not show more, that does not mean your Bescheid is right. It is more likely to miss something than to promise too much.
Limits
Where it can be wrong
- It was measured on few cases. That makes every number imprecise.
- The court cases are retold judgments, not yet descriptions by real families.
- An AI worked out from each judgment which level is right. A legal expert has not yet reviewed this.
- Some of the checker’s rules were derived from cases that later served as untouched cases. This may slightly favour the result there.
- Your own text is read by Gemini 3.8 Flash, the same model as for the current status. It has not yet been measured on the untouched cases.
What it does not do
- It does not decide on your Pflegegeld (long-term care allowance).
- It does not advise or represent you.
- The AI does not calculate. It only takes frequencies from your own words.
- It does not store your description permanently.
Is your level right?
Describe a normal day. The checker assigns your sentences and works out whether the level in the Bescheid fits.
Check my BescheidFor experts
Last measurement round: 29.09.2026.
All measurement rounds
Every scored round in which the AI reads the description, in chronological order. Nothing is overwritten.
- 26.09.2026 · Practice cases: 9 of 14Range 39 to 84 % · Model Claude Opus 5.5Erster Lauf aus freiem Text, vor der einen Änderung an der Anweisung an die KI.
- 26.09.2026 · Practice cases: 10 of 14Range 45 to 88 % · Model Claude Opus 5.5Endlauf nach dem Einfrieren von Regeln und Anweisung. Das ist die vorher festgelegte Hauptzahl.
- 26.09.2026 · Made-up cases: 7 of 7Range 65 to 100 % · Model Claude Opus 5.5Erfundene Fälle mit Antworten auf die Rückfragen, darum eine Obergrenze.
- 26.09.2026 · Untouched cases: 5 of 15Range 15 to 58 % · Model Claude Opus 5.5Unberührte Fälle, einmal gemessen, danach nichts verändert.
- 26.09.2026 · New judgments, pilot: 6 of 6Range 61 to 100 % · Model Claude Opus 5.5Pilot mit neuen Urteilen, nicht vorher festgelegt. Fast alle Fälle ohne Anspruch, sagt über Stufen darum wenig.
- 28.09.2026 · Practice cases: 10 of 14Range 45 to 88 % · Model Claude Opus 5.5 · Hours off by 17,7 on averageModellvergleich, Bezugsarm für die Qualität. Verfehlt dieselben Übungsfälle wie das Modell, das eigenen Text liest.
- 28.09.2026 · Made-up cases: 7 of 7Range 65 to 100 % · Model Claude Opus 5.5 · Hours off by 0,6 on averageModellvergleich. Erfundene Fälle mit Antworten auf die Rückfragen, darum eine Obergrenze.
- 28.09.2026 · Edge cases: 2 of 2Range 34 to 100 % · Model Claude Opus 5.5 · Hours off by 0 on averageModellvergleich, Randfälle.
- 28.09.2026 · New judgments, pilot: 6 of 6Range 61 to 100 % · Model Claude Opus 5.5 · Hours off by 5 on averageModellvergleich. Fast alle Fälle ohne Anspruch, sagt über Stufen darum wenig.
- 28.09.2026 · Practice cases: 10 of 14Range 45 to 88 % · Model Claude Sonnet 5 · Hours off by 17,3 on averageModellvergleich, bestes weiteres Modell aus der Sichtung.
- 28.09.2026 · Made-up cases: 6 of 7Range 49 to 97 % · Model Claude Sonnet 5 · Hours off by 0,7 on averageModellvergleich. Erfundene Fälle mit Antworten auf die Rückfragen, darum eine Obergrenze.
- 28.09.2026 · Edge cases: 2 of 2Range 34 to 100 % · Model Claude Sonnet 5 · Hours off by 0 on averageModellvergleich, Randfälle.
- 28.09.2026 · New judgments, pilot: 6 of 6Range 61 to 100 % · Model Claude Sonnet 5 · Hours off by 0,8 on averageModellvergleich. Fast alle Fälle ohne Anspruch, sagt über Stufen darum wenig.
- 28.09.2026 · Practice cases: 10 of 14Range 45 to 88 % · Model Gemini 3.8 Flash · Hours off by 19,4 on averageModellvergleich mit dem Modell, das eigenen Text liest, gleichauf mit Claude Opus. Gemessen über das lokale Abo-Gateway, nicht über den Weg der gehosteten Prüfung.
- 28.09.2026 · Made-up cases: 7 of 7Range 65 to 100 % · Model Gemini 3.8 Flash · Hours off by 0,6 on averageModellvergleich. Erfundene Fälle mit Antworten auf die Rückfragen, darum eine Obergrenze.
- 28.09.2026 · Edge cases: 1 of 2Range 10 to 91 % · Model Gemini 3.8 Flash · Hours off by 0 on averageModellvergleich, Randfälle. Ein Randfall blieb ohne Wertung, weil das Gateway zweimal keine Antwort lieferte; er zählt hier als nicht getroffen.
- 28.09.2026 · New judgments, pilot: 5 of 6Range 44 to 97 % · Model Gemini 3.8 Flash · Hours off by 4,2 on averageModellvergleich. Fast alle Fälle ohne Anspruch, sagt über Stufen darum wenig.
- 29.09.2026 · Practice cases: 10 of 14Range 45 to 88 % · Model Gemini 3.8 Flash · Hours off by 17,8 on averageZweite Messrunde, neu gemessen nach der Änderung der Regel zur Bindung an den Bescheid bei einer Klage (Link und Wortlaut); die Anweisung an die KI blieb gleich. Wie im Modellvergleich über das lokale Abo-Gateway gemessen, nicht über den Weg der gehosteten Prüfung.
- 29.09.2026 · Made-up cases: 7 of 7Range 65 to 100 % · Model Gemini 3.8 Flash · Hours off by 0,9 on averageZweite Messrunde. Erfundene Fälle mit Antworten auf die Rückfragen, darum eine Obergrenze.
- 29.09.2026 · Edge cases: 1 of 2Range 10 to 91 % · Model Gemini 3.8 Flash · Hours off by 0 on averageZweite Messrunde, Randfälle. Ein Randfall blieb ohne Wertung, weil das Gateway zweimal keine Antwort lieferte; er zählt hier als nicht getroffen.
- 29.09.2026 · New judgments, pilot: 5 of 6Range 44 to 97 % · Model Gemini 3.8 Flash · Hours off by 3,3 on averageZweite Messrunde. Fast alle Fälle ohne Anspruch, sagt über Stufen darum wenig.
Untouched cases
First pass: 5 of 15 cases with the reference level hit. Out of 100 similar cases, probably 15 to 58 would be right.
Of these, with a real level 4 of 13, no entitlement 1 of 1, child out of scope 0 of 1.
Missed: 10. Of these, 7 too low, 2 too high, once a child not recognised.
Correct in every run: 3 of 15.
With a follow-up question: 7 of 15.
These cases were scored once at the end, and nothing was improved on them afterwards. The result is shown here, even if it is bad.
Before the build, these cases were already part of model comparisons. 22 rules name one of them as their origin, 10 of these are in the text for the AI. This concerns X16, X19, X21, X22, X24, X25, X28, X29, X30. The case IDs themselves are not in the text for the AI. That is why we call the cases untouched since the build, not new.
Taking the level from the Bescheid
Untouched cases: on the 14 cases with a reference level, the original Bescheid was right 8 times, which is more often than Pflegegeld-Prüfer.
Practice cases: on the 14 cases with a reference level, the Bescheid was right 5 times, which is less often than Pflegegeld-Prüfer.
Across all 29 collected judgments with a Bescheid, the Bescheid was at the level of the judgment 16 times.
If its result differs from the Bescheid, that does not prove the Bescheid is wrong. It is a reason to look more closely and to get advice.
Does it show more when the Bescheid was too low?
Counted afterwards, not fixed in advance. For the next measurement, we will fix this measure in advance.
Untouched cases: When the Bescheid was below the reference level, it showed more than the Bescheid in 4 of 6 cases. When the Bescheid was right, it still showed more in 1 of 8 cases.
Practice cases: When the Bescheid was below the reference level, it showed more than the Bescheid in 3 of 3 cases. When the Bescheid was right, it still showed more in 0 of 5 cases.
Practice cases
First pass: 10 of 14 cases with the reference level hit. Out of 100 similar cases, probably 45 to 88 would be right. Target fixed in advance: 12, missed.
Of these, with a real level 6 of 9, no entitlement 1 of 1, child out of scope 3 of 4.
Correct in every run: 10 of 14.
Hit means: the majority of 3 runs hits the reference level. In addition, there is one special case in which the behaviour is scored. It counts separately; together there are 15 practice cases.
With a follow-up question: 13 of 15. Counted if it hit the level or asked the question that mattered. Target fixed in advance: 13, reached.
Missed: 4. Of these, 3 too low, none too high. Once it did not recognise a child’s case.
Calculating without AI
With the hours from the judgment, the program hits 22 of 22 cases, practice cases and made-up cases together. Target fixed in advance: 22, reached.
Untouched cases: 12 of 15. Missed: X18, X25 and X30. These judgments only give the level, no hours per task, so the program gets no input.
Made-up cases
The measurement uses 7 made-up cases. The AI reads the description, the program calculates: 7 of 7 with the expected level. Target fixed in advance: 6, reached. You can try 5 of them yourself on the checking page.
Told differently: none of 35 rewordings changed the level. This measurement was not fixed in advance.
Model for your check
Your check with your own text uses Gemini 3.8 Flash. The current status comes from this model. The details on practice cases and untouched cases in this part come from runs with Claude Opus 5.5.
A first separate run with Gemini 3.8 Flash on 26.09.2026 hit 11 of 14 practice cases. Out of 100 similar cases, probably 52 to 92 would be right. This model has not yet been measured on the untouched cases.
Model comparison
On 28.09.2026 we compared 10 AI models on the same cases: practice cases, made-up cases, edge cases and additional cases. The untouched cases were not included. On the practice cases, Gemini 3.8 Flash hit 10 of 14 and Claude Opus 5.5 10 of 14, each by the majority of runs.
Both missed the same cases: X5, X6, X12 and X15. So the gap lies more in the rules and the process than in the model. Gemini 3.8 Flash therefore remains the model for your text. Claude Opus 5.5 checks the missed cases as a second opinion, and we measure every improvement again with Gemini 3.8 Flash on the same cases.
These are few cases, and we tuned on the practice cases. So a difference between the models is not proven, in either direction.
Where the reference level comes from
The court cases come from published judgments of Austrian courts.
AI agents searched for, read and retold these judgments. For each judgment, the level that fits according to the rules was derived. Here this level is called the reference level.
In some cases it differs from the level in the judgment, for example when a child is out of scope.
The operator decided disputed questions, each time as the agents recommended. A legal expert has not yet checked the reference levels.
How we stay honest
- Practice cases and untouched cases stay separate. Nothing is improved on the untouched cases.
- What counts as a hit is fixed in advance. Changes are added as dated addenda.
- Every measurement is appended, none is overwritten. Each carries a fingerprint of the rules and of the instructions to the AI.
- Measurement runs cost no money per call, so we can measure as often as needed.
- What we counted afterwards is labelled as such.
- A check without AI runs with every change. It recalculates reports and denominators, looks for traces of the untouched cases in rules and descriptions, and compares practice cases with untouched cases. It does not measure again.
Technical terms
- Practice cases
- calibration
- Untouched cases
- holdout, already seen before the build
- Fixed in advance
- preregistration
- Range
- Wilson 95 % interval
Case by case
The practice cases one by one. Tap a case to see details.
The untouched cases one by one.
This is not legal advice. For your case, the Arbeiterkammer (Chamber of Labour) or a care counselling service can help you free of charge.