Documentation
Scheduling
The daily driver
daily.py is the incremental driver for Phases 1–3: Senate (full history, only new filings write), House (current + prior year, catches year-boundary filings), then entities.py to resolve tickers on whatever was just added.
schtasks /create /tn "Quantgress Daily" /sc daily /st 09:00 ^
/tr "cmd /c cd /d C:\path\to\Quantgress && py daily.py >> daily.log 2>&1" /f
Per-cadence phases
The remaining 12 phases run on four /etc/cron.d files, one per cadence group — not one per phase, so there are fewer files to install and fewer places to make the chmod 644 mistake. Jobs within a file are staggered so they don't all hit congress_trades.duckdb at once. Times are server-local. Source: deploy/cron.d/ in the repo.
quantgress-daily — every day
| Time | Script | Phase |
|---|---|---|
| 06:00 | daily.py | 1–3 (Senate + House scrapers, ticker resolution) |
| 06:15 | scrape_short_volume.py | 11 (FINRA short volume) |
| 06:30 | scrape_insiders.py | 9 (SEC Form 4 insiders) |
quantgress-weekly — Sundays
A cheap safety margin, not a real upstream deadline for either source.
| Time | Script | Phase |
|---|---|---|
| Sun 07:00 | scrape_contracts.py | 7 (government contracts) |
| Sun 07:15 | scrape_patents.py | 12 (USPTO patents) |
quantgress-quarterly — Jan / Apr / Jul / Oct 1st
The only group whose cadence is genuinely dictated by the source — 13F and lobbying filings are quarterly upstream.
| Time | Script | Phase |
|---|---|---|
| 08:00 | scrape_13f.py | 10 (13F institutional holdings) |
| 08:15 | scrape_lobbying.py | 6 (LDA lobbying filings) |
quantgress-annual — Jan 2nd
Runs early January, deliberately after year-end filings settle.
| Time | Script | Phase |
|---|---|---|
| 09:00 | scrape_donors.py | 13 (OpenFEC donations) |
| 09:15 | scrape_execcomp.py | 16 (executive comp) |
| 09:30 | scrape_senate_annual.py | 18 (Senate annual financial disclosures) |
Deliberately not on any cron schedule
Phase 17 — scrape_trump.py. The ProPublica DocumentCloud mirror this phase reads has no fixed publication cadence to key a cron line off of. Run manually (py scrape_trump.py) until a pattern emerges.
Phase 14 — scrape_pageviews.py. Requires an explicit --article watchlist; no ticker→Wikipedia-title mapping exists yet, so an unattended call fails on a usage error every time. Was in quantgress-weekly originally, pulled after a live run hit exactly that failure. Run by hand with --article until a watchlist exists.
Install
sudo cp deploy/cron.d/quantgress-* /etc/cron.d/
sudo chmod 644 /etc/cron.d/quantgress-*
mkdir -p /home/ubuntu/quantgress/logs
Cron silently ignores any /etc/cron.d file with group/other write permission — no error, the job just never runs. chmod 644 after every copy, not just the first.