QA & Review
The review workflow for checking scraper output against the rubric before merging.
QA & Review
Reviewing a scraper PR comes down to one question: is the output correct? The QA workflow exists to answer that quickly and consistently.
The review process
Pull the PR and run the spider locally:
scrapy crawl {spider} -O data/scrapers/{spider}.jsonOpen the viewer at http://localhost:3000/scrapers/<spider-name> and scan
the output - or paste the JSON into the
rubric report.
Apply the rubric - check each field against the QA rubric. The rubric is designed to be applied field by field, in order.
Check the schema - make sure every required field is present and matches the Schema reference.
Leave review comments referencing specific rubric rules, e.g. "field
links is out of priority order - see
/docs/qa/schema#links".
The fastest way to review is the viewer in one tab and the rubric in another. The viewer shows you the data; the rubric tells you what to check.
What to block on
- Missing required fields (
id,title,start,status,source). - Datetimes that carry timezone info (
start/endmust be naive). - Duplicate records (same start time + similar title).
- Status inconsistent with
starttime (e.g.passedfor a future meeting). - Status or classification values outside the schema vocabularies.
- Links out of priority order.
What to comment on (not block)
- Title formatting (should include the body name).
- Description quality (should be useful, not just a copy of the title).
time_noteswhen location is empty (should explain why).
Where this fits
The rubric answers "is this output correct" for one PR. For the whole multi-team PR lifecycle, see Review process. For the recurring findings PDW catalogued across 23 real PRs, see Review issues.
Last updated on