Building Scrapers

Building Scrapers

The development workflow for writing city-meeting scrapers, from source analysis to PR.

Repo
city-scrapers (core + consumer repos)city-scrapers-core and the per-city repos like city-scrapers-fortx

Building scrapers

City Scrapers uses Scrapy to extract public meeting information from government websites. Each spider produces a JSON array of meeting records that flows into Documenters.org - and that you can inspect in the Meetings Viewer while you develop.

The workflow

Scope the target - decide whether the agency's site is worth scraping at all, and whether anyone has already scoped it. See Scoping a target.

Analyze the source - open the agency's meeting page in DevTools, check the Network tab for a JSON endpoint, and pick a primary URL. See Source analysis.

Pick a spider pattern - LegistarSpider for Legistar sites, a spider factory for many similar agencies, or a singular CityScrapersSpider for everything else. See Spider patterns.

Write the spider - implement parse and the _parse_* helpers, following the schema field by field. See Writing a spider.

Run and inspect - scrapy crawl <name> -O <name>.json, then load the file in the viewer and apply the QA rubric.

Test and submit - write pytest tests against saved fixtures, run flake8/black/isort, open a PR. See Testing and Contributing.

The tight loop is steps 4 and 5: run the spider, refresh the viewer, fix, repeat. The viewer exists to make that loop fast.

Common pitfalls

  • Lowercase -o: silently appends to an existing file, duplicating records. Always use uppercase -O.
  • Timezone-aware datetimes: start and end must be naive. The spider's timezone class attribute declares the local timezone; never embed tzinfo in the datetime itself.
  • Missing time defaults: if a source has a date but no time, default to midnight and say so in time_notes.
  • Links out of order: the priority is agenda, minutes, video, then other attachments - not agenda, video, minutes.
  • async def start(): consumer repos pin Scrapy 2.11.2, which does not recognize the newer start() coroutine. Use start_requests().

Section map

Last updated on