Auditing AI-generated trial summaries
A lay summary can sound clear while overstating a benefit, changing a number or omitting a clinically important caveat. I built a generator and a separate auditor to compare summaries with a structured evidence table and classify those errors.
Two published trials, PARAGON-HF and SELECT, have field-traceable evidence tables and annotated base summaries. In PARAGON-HF, the primary outcome did not meet conventional statistical significance; exploratory findings must not be presented as confirmed efficacy.
The project defines 17 error codes (six commission, eleven omission) and has a four-case development panel: three cases with injected errors and one reviewed negative. The final panel has not been run, the reference standard has not been annotated, and no performance metric has been computed.
Evaluation workflow
- Extract published study findings into a field-traceable evidence table.
- Generate a lay summary from the table and inject a known error into selected development examples.
- Give a separate auditor the summary and the same table, without the source document or the inserted-error label.
- Compare coded findings with prespecified targets; a human-annotated reference standard is still needed for the final evaluation.