Backtesting invoice classification on the billing@alpha.school mailbox against QuickBooks · Initiative B · September 28, 2026
Three results matter most for building this:
Every invoice Alpha School pays reaches billing@alpha.school by email, and every bill it books lands in one of ~73 QuickBooks companies, one per legal entity (campus). The task: given an email as it arrives, predict the company its bill will be booked to — or say it can't tell.
All Mail from 1 January to
27 September 2026, with headers, bodies, attachments, Gmail thread ids and the
AP team's labels. Read/unread state was never touched.A prediction may only use what existed when the email arrived. That rule was enforced mechanically, and an audit partway through found and fixed two violations:
Three families, all observable at arrival:
Eleven approaches, from trivial to stacked, all scored on the same September emails:
| Approach | What it is | Learns from labels? |
|---|---|---|
| Majority | always the most common company | — |
| Sender history | the company most of this sender's earlier bills went to | no (running count) |
| Rules | a Bill-To ZIP, any single ZIP, or one campus name in the text | no |
| Jev, 75 labels | TypeSafe's Jev Choice model over all 73 companies + 2 catch-alls |
no (zero-shot) |
| Jev, 16 labels | the same, restricted to the 15 candidate companies + "not campus-specific" | no |
| LLM zero-shot | GLM 5.3 Flash asked directly, may answer "unknown" | no |
| Logistic regression | text + tabular features, one-hot per company | yes |
| Gradient boosting | tabular features only, one-hot per company | yes |
| Candidate ranker, no LLM | the stacked model below without the LLM features | yes |
| Candidate ranker | scores every (email, candidate company) pair on relationship features and picks the top | yes |
| Routing ensemble | gradient boosting decides, unless the LLM names a specific non-HQ campus with confidence ≥ 0.9 (cutoff chosen on August) | yes |
The candidate ranker is the stacked model. Rather than one-hot features per company, each row is an (email, company) pair described by features such as this sender's past share of bills to this company, the LLM's confidence for this company and the out-of-fold text probability for this company. The features mean the same thing for every company, so one model covers all of them — including campuses that have few or no past bills, which is the situation every time Alpha opens one.
Every method may abstain, so each is reported at three operating points: always answering (abstentions count as wrong), and on the emails it answers with confidence ≥ 0.7 and ≥ 0.9, with the share it answers (coverage).
The table scores every method on the same 958 September emails (15 companies; Austin K-8 is 34% of them, so "always Austin" scores 34%). The chart shows the trade-off each method offers between answering more emails and being right more often.
| Method | Always answering | Accuracy ≥ 0.7 | Coverage | Accuracy ≥ 0.9 | Coverage |
|---|---|---|---|---|---|
| Majority company | 34.4% [32%–38%] | 34.4% [32%–38%] | 100% | 34.4% | 100% |
| Sender history | 41.4% [38%–44%] | 56.6% [53%–61%] | 69% | 90.1% | 17% |
| Rules (ZIP / campus name) | 30.0% [27%–33%] | 50.8% [47%–55%] | 59% | 50.8% | 59% |
| Jev, 75 labels | 34.1% [31%–37%] | 56.2% [52%–61%] | 51% | 61.5% | 39% |
| Jev, 16 labels | 39.4% [36%–43%] | 79.8% [76%–84%] | 46% | 86.6% | 35% |
| LLM zero-shot (GLM 5.3 Flash) | 44.4% [41%–48%] | 84.5% [81%–88%] | 49% | 88.7% | 38% |
| Logistic regression | 79.1% [77%–82%] | 90.1% [88%–92%] | 76% | 94.3% | 59% |
| Gradient boosting (per company) | 80.6% [78%–83%] | 89.0% [87%–91%] | 84% | 92.8% | 68% |
| Candidate ranker, no LLM | 71.0% [68%–74%] | 80.0% [77%–83%] | 82% | 85.6% | 61% |
| Candidate ranker | 74.5% [72%–77%] | 82.0% [79%–84%] | 79% | 88.0% | 59% |
| Routing ensemble (boosting + LLM override) | 84.0% [82%–86%] | 89.3% [87%–91%] | 90% | 91.4% | 79% |
Brackets: bootstrap 95% confidence intervals. “Always answering” counts abstentions as wrong. Coverage = share of messages the method answers at that confidence.
| Slice | n | Jev, 16 labels | LLM zero-shot (GLM 5.3 Flash) | Logistic regression | Gradient boosting (per company) | Routing ensemble (boosting + LLM override) |
|---|---|---|---|---|---|---|
| New invoice | 402 | 61% | 68% | 75% | 76% | 84% |
| Follow-up | 556 | 24% | 27% | 82% | 84% | 84% |
| Employee forward | 384 | 38% | 41% | 82% | 84% | 84% |
| Via invoicing platform | 192 | 51% | 59% | 76% | 80% | 87% |
| Vendor direct | 143 | 63% | 69% | 82% | 76% | 93% |
| Employee thread | 143 | 16% | 20% | 67% | 68% | 66% |
| Vendor reply | 92 | 20% | 25% | 89% | 93% | 92% |
| All other companies | 628 | 50% | 57% | 73% | 73% | 82% |
| Austin K-8 | 330 | 19% | 20% | 90% | 95% | 88% |
Accuracy always answering, per slice. “New invoice”: the bill did not exist in QuickBooks when the email arrived.
| Company | n | Jev, 16 labels | LLM zero-shot (GLM 5.3 Flash) | Logistic regression | Gradient boosting (per company) | Routing ensemble (boosting + LLM override) |
|---|---|---|---|---|---|---|
| Alpha School 78746, LLC (Austin K-8) | 330 | 19% | 20% | 90% | 95% | 88% |
| Alpha School 33155, LLC | 153 | 40% | 43% | 93% | 97% | 97% |
| Alpha School 60601, LLC (Chicago) | 91 | 65% | 65% | 37% | 31% | 65% |
| Alpha School 78701, LLC (Alpha High School) | 62 | 32% | 48% | 92% | 89% | 90% |
| Sports Academy 75010, LLC (CSA) | 56 | 43% | 64% | 89% | 89% | 89% |
| Alpha School 00692, LLC (Dorado) | 48 | 67% | 69% | 79% | 92% | 92% |
| Alpha School 94123, LLC (San Francisco) | 44 | 55% | 45% | 82% | 61% | 64% |
| Alpha School 92610, LLC (Orange County) | 38 | 66% | 76% | 82% | 82% | 89% |
| Sports Academy 78734, LLC (TSA) | 36 | 44% | 75% | 67% | 75% | 75% |
| Alpha Early Center 10038, LLC (Alpha Early Center Fin District) | 28 | 39% | 46% | 0% | 0% | 39% |
| Alpha School 94306, LLC (Palo Alto) | 21 | 71% | 71% | 90% | 81% | 81% |
| Alpha School 78521, LLC (Brownsville) | 19 | 32% | 32% | 74% | 100% | 100% |
| gt.school 78626, LLC | 13 | 92% | 100% | 100% | 100% | 100% |
| Alpha School 94965, LLC (Alpha Marin) | 11 | 73% | 73% | 18% | 0% | 73% |
| Alpha Anywhere Center 10504, LLC (Alpha Armonk) | 8 | 25% | 62% | 0% | 0% | 12% |
| Method | Always answering | Accuracy ≥ 0.7 | Coverage |
|---|---|---|---|
| Majority company | 48.2% | 48.2% | 100% |
| Sender history | 36.1% | 54.3% | 63% |
| Rules (ZIP / campus name) | 20.6% | 36.9% | 56% |
| LLM zero-shot (GLM 5.3 Flash) | 36.4% | 77.1% | 43% |
| Logistic regression | 74.1% | 82.8% | 80% |
| Gradient boosting (per company) | 73.4% | 79.4% | 86% |
| Candidate ranker, no LLM | 75.8% | 78.3% | 90% |
| Candidate ranker | 76.8% | 79.7% | 89% |
| Routing ensemble (boosting + LLM override) | 77.3% | 82.5% | 89% |
Fit on 4,977 linked messages (January–July), scored on 772 August messages. All design choices were made here.
To see which inputs carry the signal, the candidate ranker was refit on August with one feature group removed at a time. Every single group moves accuracy by only one to three points: the information is spread across overlapping signals (the vendor's name appears in the text, in the LLM's extraction and in the vendor-history features), so no one group is indispensable. Removing the two content models together (LLM and text) costs the most. The per-slice results are the clearer evidence of what matters: follow-ups and Austin — where history decides — are where the zero-shot methods collapse and the learned ones do not.
The learning curve (below) was run on the ranker. It does not rise with more data: 10% of the training emails does about as well as all of them. The limit is ambiguity (the HQ address, shared vendors), not volume.
| Feature group removed | Change (points) |
|---|---|
| LLM + text model | -3.1 |
| Text model | -2.5 |
| All context | -2.2 |
| LLM extraction | -1.0 |
| Company prior | -1.0 |
| Thread | -1.0 |
| Origin | -0.8 |
| ZIP / campus-name matches | -0.5 |
| Vendor history | -0.3 |
| Sender-domain history | +0.1 |
Recommendation. Build the production classifier as the routing ensemble: gradient boosting over the email, QuickBooks-history and LLM-extracted features, with the LLM allowed to override when it names a specific non-HQ campus at ≥ 0.9 confidence. Auto-assign when the combined confidence is ≥ 0.9 (79% of emails at 91% accuracy in this backtest) and send the rest to a person with the top suggestions. Do not use rules or Jev as the classifier.
Next steps, in order of expected gain:
Everything is in the AlphaSchoolAttgTriage repository; derived data lives in
the project's private S3 bucket (scripts/data_sync.py pull). In order:
export_alpha_mail.py --mailbox "[Gmail]/All Mail" --since 2026-01-01 --out mail_archive — read-only mailbox exportscripts/qbo_discover.py — every authorized QuickBooks company: accounts, classes, vendors, billsscripts/qbo_match_mail.py --mail-dir mail_archive — link emails to bills (the ground truth)scripts/llm_extract.py --model fw-fireworks-trilogy-glm-5p3-flash — LLM fields per linked emailscripts/build_backtest.py — one row per email: as-of features, outcomes, splitscripts/backtest_ml.py, scripts/backtest_rank.py (and --drop GROUP, --fit-frac F) — learned models, ablations, learning curveclassify_locations_jev.py (JEV_CHOICES=active for the 16-label run) — Jevscripts/run_experiments.py — every method scored on the same rowsscripts/build_report.py — this documentFull method and protocol notes: docs/approaches.md; history of experiments: docs/experiments.md.