Key takeaways
- Twenty cases per jurisdiction and six model interfaces define the benchmark.
- Repeated cases, models and materials create non-independent observations.
- Reasoning verification and a human-lawyer baseline are not completed results.
What was tested & why
Limitations
Between-stage comparisons can be descriptive without being controlled. Stated confidence is elicited from models; the sample is imbalanced toward Dismissed appeals in the Nigerian subset. Unverified reasoning and exploratory tests should not be promoted into established findings.
Source: Thesis §10.10 · source-reported unless otherwise noted.