← back to all models ← RevenueBench home
aion-labs

aion-labs/aion-3.5-mini

Mean 0.926 · 28/36 perfect tests · $0.72 total · 36/40 tests completed — remaining tests discarded from study (endpoint timeouts after dedicated retries) · modalities in: text · out: text · each card: the prompt → the correct answer (gold) → this model's actual answer

What do the modalities mean?

Modalities in means what you can send this model: text, images, files, video, or audio. Modalities out means what it can send back. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what this model can handle in real work.

data-crm 4 tests
owner-resolutionPASS 1.00
Plain-English question: Every open deal has a blank owner name — resolve each one to a real person using the owners table.
Prompt — what the model saw
The deal extract below has an owner name field that is blank on nearly every row and an owner id field that is populated. Using the owners table, resolve each open deal to an owner name. Report: number of deals resolved, deals whose owner id has no match in the owners table (list them), the owner ids that map to archived/deactivated owners (list their names), and the total pipeline amount per resolved owner. Do not guess a name for any unmatched id.
cost $0.0206 · 98s · in 12,177 / out 6,038 tok
✓ resolved_count✓ unmatched_none✓ deactivated_flagged✓ jake_amount✓ justin_amount✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: All 156 deals resolve to six named owners; jake_amount maps to Bryce Harmon ($1,054,144) and justin_amount to Alex Franklin ($624,310).
The answer in one line: jake_amount = 1,054,144; justin_amount = 624,310; resolved = 156
jake_amount1,054,144
justin_amount624,310
resolved156
deactivated
per_ownerBryce Harmon=1,054,144; Alex Franklin=624,310; Dana Mercer=341,195; Cole Ingram=288,161.43; Farid Osman=4,134; Elena Sinclair=2,100
graded checksresolved_count · unmatched_none · deactivated_flagged · jake_amount · justin_amount · no_invented

Model's actual answer

show the model's full answer
Owner resolution for the open-deal extract (deals_open.csv joined to owners.csv on owner_id):

1) Deals resolved to an owner name
All 156 open deals resolved. Every owner_id present in the deal file has a match in owners.csv.

2) Deals whose owner_id has no match in owners.csv
None. Owner ids appearing in the deals (119337721, 84342457, 83155923, 83155924, 716654662, 701163055) all exist in the owners table.

3) Owner ids mapping to archived/deactivated owners
None of the deal owner ids map to archived owners. Two owners in owners.csv are archived=true — Gavin Porter (1520255671) and Hugo Lindqvist (77260721) — but neither id appears on any open deal.

4) Total pipeline amount per resolved owner (sum of amount, all stages/forecast categories)

- Bryce Harmon (119337721) — 35 deals = 1,054,144
  (24000+19656+13500+7000+2520+240000+99000+72000+70000+63600+45000+1+21000+23400+13680+5502+8160+1+11400+1+36000+31500+6000+10800+30275+17400+12600+18000+37440+18828+2880+36000+20880+10920+25200)
- Alex Franklin (84342457) — 67 deals = 624,310
  (14850+13770+11200+9000+6360+5400+3240+2484+1920+1080+7200+19000+2880+1400+4800+1632+10000+9300+2700+2160+1800+3600+3840+15000+1968+4000+3600+4800+3120+2520+9000+2400+62000+5400+5100+16700+4400+1620+2600+7200+18000+17000+8316+8100+18000+12600+24000+15000+9000+7200+3780+16200+7200+4680+1800+18000+2730+2400+3060+18000+12000+1800+4400+31200+7200+1600+60000)
- Dana Mercer (83155923) — 24 deals = 341,195
  (11250+10500+9000+9000+5400+4800+4600+1920+15000+4200+18900+27000+43875+20000+60000+8100+16250+3150+5000+2100+23400+5400+7350+25000)
- Cole Ingram (83155924) — 22 deals = 288,161.43
  (2760+1330+7225.4+3360+5616+2700+7781.2+6947.5+40000+12168+2480.4+11116+11193+1875+31750+32175+4779.88+3334.8+58529.25+4140+18000+18900)
- Farid Osman (716654662) — 7 deals = 4,133
  (1249+1+1+1+2880+1+1)
- Elena Sinclair (701163055) — 1 deal = 2,100
  (Deal-57FF13 = 2100)

Deal count check: 35+67+24+22+7+1 = 156, matching the extract.
Grand total pipeline: 1,054,144 + 624,310 + 341,195 + 288,161.43 + 4,133 + 2,100 = 2,314,043.43
gong-hubspot-join-integrity0.75
Plain-English question: How many call records point at deal ids that do not exist in the open pipeline?
Prompt — what the model saw
Check the integrity of the call-to-deal mapping below. The gong table lists calls in the last 90 days keyed to HubSpot deal ids; deals_open lists the current open default-pipeline deals. Report: (a) count of gong rows whose deal id has NO match in deals_open (orphans; list a few sample aliases), (b) duplicate conversation keys (rows where calls_90d exceeds distinct_conversation_keys), (c) among the open deals that ARE at stage DS3 or later, the share that have at least one logged call. Give sample aliases for the orphan finding, all taken from the data. Note: gong_owner_id is a Gong user id, not a HubSpot owner id.
cost $0.0487 · 390s · in 15,817 / out 24,292 tok
✓ orphan_count✗ dup_keys_0✓ ds3_coverage✓ sample_ids_present

✓ Correct Answer (ground truth)

Plain-English answer: 30 Gong call rows are orphans keyed to dead deal ids, none are duplicates, and DS3 call coverage is 29.4%.
The answer in one line: orphans = 30; ds3_cov_pct = 29.4; dup = 0
orphans30
ds3_cov_pct29.4
dup0
orphan_aliasesDeal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42, Deal-36EA09
graded checksorphan_count · dup_keys_0 · ds3_coverage · sample_ids_present

Model's actual answer

show the model's full answer
Call-to-deal mapping integrity check (gong_calls_by_deal_90d.csv vs deals_open.csv)

(a) Orphan gong rows (deal id not in deals_open)
- 30 of 67 gong rows have no match in deals_open: 30/67 = 0.448 → 44.8%.
- Sample orphan aliases (all from the data): Deal-8FA85D (60251290957, 46 calls), Deal-8FC3F9 (60251649055, 24), Deal-3B7945 (60251639682, 21), Deal-42B265 (61227242540, 21), Deal-9CCC42 (61430316324, 17), Deal-36EA09 (60251082126, 17).
- Pattern note: all 30 orphans sit in the 60251xxxxx/6025xxxxx id range block of the file and all show calls_with_ds2_date = 0 and calls_with_sqo_date = 0; several also carry last_call dates after 2026-09-24 (e.g., Deal-8FA85D last_call 2026-11-05, Deal-42B265 2026-11-25), which is inconsistent with a "last 90 days" window as of late September 2026.

(b) Duplicate conversation keys
- 0 rows. Every gong row has calls_90d exactly equal to distinct_conversation_keys (e.g., Deal-C6FE92 11=11, Deal-8FA85D 46=46, Deal-3B6668 4=4). No row exists where calls_90d > distinct_conversation_keys.

(c) Call coverage among open deals at DS3 or later
- Stage mix of the 156 open deals: 32 DS1 + 39 DS2 + 61 DS3 + 14 DS4 + 10 DS5 (32+39+61+14+10 = 156).
- DS3-or-later population: 61 + 14 + 10 = 85 deals.
- Of those 85, deals with at least one logged call (i.e., appearing in the gong extract; every gong row has calls_90d ≥ 3):
  - DS4/DS5: 13 of 24 have calls (e.g., Deal-25F752, Deal-E53952, Deal-5408B0, Deal-D348E1, Deal-C26D20, Deal-547B2B, Deal-944310, Deal-B7EBD1, Deal-3974EB, Deal-A2B47C, Deal-C61CF7, Deal-584EE5, Deal-C6D97A).
  - DS3: 12 of 61 have calls (Deal-C6FE92, Deal-40522D, Deal-333EBB, Deal-60C2C2, Deal-1E2498, Deal-D9A12F, Deal-215CCA, Deal-5EED42, Deal-DAF1D9, Deal-8952F0, Deal-BA3DDC, Deal-F9A3C1).
- Share with at least one logged call: (13 + 12) / (24 + 61) = 25/85 = 0.294 → 29.4%.

Caveats
- (c) measures calls inside the 90-day gong extract only; a DS3+ deal with calls older than the window would be undercounted, and that cannot be verified from the data provided.
- Per your note, gong_owner_id was treated as a Gong user id and not joined to deals_open.owner_id — the id spaces differ (e.g., gong 4729257026550273695 vs HubSpot 119337721), consistent with that.
snowflake-sql-trapped-catalogPASS 1.00
Plain-English question: Write one Snowflake query for first-month usage per customer using only the provided, trap-laden catalog.
Prompt — what the model saw
Using ONLY the catalog provided, write one Snowflake SQL query returning, per customer company, for its first calendar month as a customer: unique givers, recognition count, successful redemption count. Read the catalog notes carefully: some tables are stale or unpopulated and must be avoided, and one documented filter (excluding deleted givers) must NOT be applied to historical giving counts. Put a comment above each FROM or JOIN explaining the table choice. Return only fenced SQL.
cost $0.0158 · 206s · in 0 / out 8,713 tok
✓ avoids_stale_deals✓ avoids_unpopulated_live✓ uses_cohort_summary✓ succeeded_filter✓ no_giver_deletion_filter✓ returns_sql

✓ Correct Answer (ground truth)

Plain-English answer: A single correct SQL query that uses the right catalog tables and avoids the trap columns — graded on the exact tables and columns.
The answer in one line: use_table = PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
avoid_stalePRODUCTION.HUBSPOT.DEALS (stale)
avoid_unpopulatedV2_LIVE.OBJECTS_DEALS (unpopulated)
forbidden_filterdeleted-giver exclusion must NOT be applied
outputone fenced SELECT with per-FROM/JOIN table-choice comment
required_filtersucceeded redemptions only
use_tablePRODUCTION.PLG.COMPANY_COHORT_SUMMARY
graded checksavoids_stale_deals · avoids_unpopulated_live · uses_cohort_summary · succeeded_filter · no_giver_deletion_filter · returns_sql

Model's actual answer

show the model's full answer
```sql
-- ============================================================================
-- Per customer company, first calendar month as a customer:
--   unique givers | recognition count | successful redemption count
-- ============================================================================
-- EXPLICIT DATA GAPS (stated per instructions; nothing below is invented):
--
-- (1) NO recognition/giving events table exists in this catalog. The only
--     giving data available is the pre-aggregated cohort columns M1_USERS /
--     M1_GIVES / M1_REDEMPTIONS in PRODUCTION.PLG.COMPANY_COHORT_SUMMARY.
--     Therefore:
--       * "Unique givers" is reported via M1_USERS. The catalog documents
--         this column as USERS, not givers specifically; a true giver-level
--         count is not obtainable from this catalog.
--       * "Recognition count" is reported via M1_GIVES.
--     The catalog does not document whether M1 is the first CALENDAR month
--     or a rolling ~30-day window from the cohort anchor date. If M1 is a
--     rolling window, these figures will not exactly equal a calendar-month
--     cut. Flagged as an assumption, not stated as fact.
--
-- (2) PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 is the
--     documented source for redemption counts (STATE = 'succeeded' only),
--     but the catalog excerpt documents NO columns for it -- no company
--     key, no event timestamp -- and no join path to the cohort table.
--     Querying it would require inventing column names, so it is NOT
--     referenced here. Successful redemptions are taken from M1_REDEMPTIONS
--     (per the catalog's business rules, redemptions = STATE 'succeeded'
--     only, which these cohort aggregates reflect).
--
-- (3) Deleted-giver exclusion filter: NOT applied, per the documented rule
--     that it understates historical giving counts. (No giving events table
--     exists in this catalog, so no such filter could be applied in any case.)
--
-- (4) Scope: COMPANY_COHORT_SUMMARY covers SELF-SERVE companies only.
--     Sales-led customers are not represented anywhere in this catalog and
--     are therefore out of scope.
--
-- (5) "First calendar month as a customer": the catalog offers SIGNUP_DATE,
--     ACTIVATED_DATE, and FIRST_SUB_PAYMENT_DATE without defining
--     "customer." SIGNUP_DATE is used here for consistency with the M1_*
--     cohort aggregates (presumed signup-anchored; actual anchor is
--     undocumented). If the business definition is "paying customer," swap
--     the anchor to FIRST_SUB_PAYMENT_DATE.
--
-- (6) The cohort table's company identifier column is not named in the
--     catalog excerpt; COMPANY_ID is used as a placeholder key and must be
--     verified against the real table before running.
--
-- Tables deliberately avoided per catalog notes:
--   HUBSPOT_HUB_1973303.V2_LIVE.OBJECTS_DEALS  (exists but UNPOPULATED)
--   PRODUCTION.HUBSPOT.DEALS                   (STALE -- last sync 2023-03)
--   PRODUCTION.HUBSPOT.GONG_HUBSPOT_MAP_FAST   (RETIRED pattern)
-- ============================================================================

SELECT
    c.COMPANY_ID                         AS customer_company,              -- placeholder key; see gap (6)
    DATE_TRUNC('month', c.SIGNUP_DATE)   AS first_customer_month,          -- see gap (5)
    c.M1_USERS                           AS unique_givers,                 -- approximated from users; see gap (1)
    c.M1_GIVES                           AS recognition_count,             -- see gap (1)
    c.M1_REDEMPTIONS                     AS successful_redemption_count    -- succeeded-only per business rule; see gap (2)
-- Table choice: PRODUCTION.PLG.COMPANY_COHORT_SUMMARY is the only table in
-- the catalog with per-company first-month giving and redemption figures
-- (one row per self-serve company). It is current (no staleness note), and
-- the stale / unpopulated / retired tables above are avoided. No JOIN is
-- possible or needed: no other catalog table provides giving events, and
-- REDEMPTION_RECORDS_V2 lacks documented join keys (see gap 2).
FROM PRODUCTION.PLG.COMPANY_COHORT_SUMMARY AS c;
```
crm-completeness-enrichmentPASS 1.00
Plain-English question: Audit the CRM for missing and invalid fields, and propose the fixes.
Prompt — what the model saw
Audit the CRM extract for completeness and propose fixes. Deals need owner, stage, amount, close date, why-buys; companies need industry, employee count, HQ country; contacts need email, title, persona. Report completeness percent per field. Identify duplicate company clusters (name variants, shared domain) with a survivor per cluster. List invalid emails and domain mismatches. Fill missing company fields only where the enrichment export has a matching row; where CRM and enrichment disagree list both and recommend a source. Never invent a value. End with the 10 fixes with the most pipeline amount at stake.
cost $0.0232 · 226s · in 1,643 / out 13,139 tok
✓ dup_clusters✓ invalid_emails✓ survivor_language✓ completeness_pct✓ no_invent

✓ Correct Answer (ground truth)

Plain-English answer: 9 companies missing employee count, 6 missing HQ country, 4 invalid emails, 1 domain mismatch, 2 duplicate clusters, and 25 enrichment matches.
The answer in one line: n_contacts = 52; n_companies = 34; enrichment_matches = 25; invalid_email_count = 4; n_dup_clusters = 2; domain_mismatch_count = 1 (+4 more below)
n_contacts52
domain_mismatch_count1
enrichment_matches25
invalid_email_count4
n_companies34
n_dup_clusters2
company_missingindustry=0; employee_count=9; hq_country=6
domain_mismatchesCT-0011
duplicate_clusters['C-0A092931', 'C-0A092932'], ['C-0A092933', 'C-0A092934']
invalid_emailsCT-0010, CT-0080, CT-0081, CT-0192
graded checksdup_clusters · invalid_emails · survivor_language · completeness_pct · no_invent

Model's actual answer

show the model's full answer
CRM EXTRACT AUDIT — COMPLETENESS, DUPLICATES, EMAILS, ENRICHMENT FILLS, TOP 10 FIXES

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
1. SCOPE GAP (stated first, per "say so explicitly")
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
No deals file was provided in the extract (only companies.csv, contacts.csv, zoominfo_enrichment.csv). Therefore:
- Deal-field completeness (owner, stage, amount, close date, why-buys) CANNOT be computed — 0 deals audited.
- "Pipeline amount at stake" CANNOT be computed from deal amounts. The Top-10 ranking below uses a stated proxy: number of affected company records (deals attach to companies) and contacts. Re-rank by open pipeline per company once deals.csv is supplied.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
2. COMPLETENESS BY FIELD (showing arithmetic)
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
DEALS — file absent: 0% auditable on all five fields (owner, stage, amount, close date, why-buys).

COMPANIES (n = 34 rows)
- domain: 34/34 = 100%
- industry: 34/34 = 100% populated — but non-standardized: "tech" (C-66D1FC, C-44EA29, C-60C75F), "Tech " with trailing space (C-425E2A, C-BA969B, C-93C8BF, C-C9BB20), "health care" (C-7BBDFA, C-50D386), "SaaS" (C-0A092933), vs "Technology"/"Healthcare"/"Retail"/"Finance"/"Manufacturing" elsewhere.
- employee_count: blanks at C-EC3025, C-96039F, C-44EA29, C-D04904, C-B23205, C-60C75F, C-7BBDFA, C-50D386, C-93C8BF = 9 blanks → 25/34 = 73.5%
- hq_country: blanks at C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB = 6 blanks → 28/34 = 82.4%. Format variants: US (10), USA (5), United States (2), UK (3), Canada (8).

CONTACTS (n = 52 rows)
- email present: 52/52 = 100%; WELL-FORMED: 48/52 = 92.3% (4 malformed, see §4)
- title: blanks at CT-0000, CT-0022, CT-0072, CT-0080, CT-0081, CT-0092, CT-0120, CT-0121, CT-0122, CT-0132, CT-0141, CT-0162, CT-0170 = 13 blanks → 39/52 = 75.0%
- persona: blanks at CT-0000, CT-0022, CT-0041, CT-0060, CT-0070, CT-0081, CT-0082, CT-0092, CT-0110, CT-0132, CT-0162, CT-0171, CT-0172, CT-0180, CT-0181 = 15 blanks → 37/52 = 71.2%
- Unique contacts with ANY gap: 6 both-blank + 7 title-only + 9 persona-only = 22 of 52 (42.3%).

Also: 14 of 34 companies (41.2%) have zero contacts — C-2C60E5, C-B97B4E, C-2D7423, C-7BBDFA, C-50D386, C-BA969B, C-332637, C-93C8BF, C-EE9FFB, C-C9BB20, C-0A092931, C-0A092932, C-0A092933, C-0A092934.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
3. DUPLICATE COMPANY CLUSTERS
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Note: company names are not in the extract (aliases are opaque IDs), so clustering is by shared domain only. Name-variant detection is not possible with the data given.

Cluster 1 — domain acme-corp.com
- C-0A092931 (Technology, 500, US)
- C-0A092932 (tech, 510, USA)
- Survivor: C-0A092931 (earliest record). Merge C-0A092932 in. Employee count conflict 500 vs 510 is unresolved — no enrichment row for acme-corp.com to arbitrate; do not pick a value without a source.

Cluster 2 — domain globex.io
- C-0A092933 (SaaS, 200, US)
- C-0A092934 (Technology, 200, US)
- Survivor: C-0A092933 (earliest record). Industry conflict SaaS vs Technology — see §5 recommendation.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
4. INVALID EMAILS AND DOMAIN MISMATCHES
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Malformed (no domain after @) — 4:
- CT-0010 (C-66D1FC): user0@
- CT-0080 (C-92D97D): user0@
- CT-0081 (C-92D97D): user1@
- CT-0192 (C-425E2A): user2@
Repair candidates from the contacts' own domain column (verify deliverability before committing; this is reconstruction from in-file data, not an invented external value): user0@66d1fc.com, user0@92d97d.com, user1@92d97d.com, user2@425e2a.com.

Domain mismatch — 1:
- CT-0011 (C-66D1FC): user1@other-domain.com vs company domain 66d1fc.com (contact.domain column also says 66d1fc.com). Verify whether this is a legitimate alternate/personal domain or a data-entry error.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
5. ENRICHMENT FILLS AND CRM-vs-ENRICHMENT CONFLICTS
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Enrichment covers 25 domains; 7 CRM domains have NO enrichment row: acme-corp.com, globex.io, ba969b.com (C-BA969B), 332637.com (C-332637), 93c8bf.com (C-93C8BF), ee9ffb.com (C-EE9FFB), c9bb20.com (C-C9BB20).

FILLS (only where enrichment has a matching row; all ZI employee_count = 400):
- C-EC3025: employee_count blank → 400
- C-96039F: blank → 400
- C-44EA29: blank → 400
- C-D04904: blank → 400
- C-B23205: blank → 400
- C-60C75F: blank → 400
- C-7BBDFA: blank → 400
- C-50D386: blank → 400
Total: 8 fills → employee_count completeness would rise 73.5% → 97.1% (33/34). C-93C8BF stays blank (no ZI row).

hq_country: ZERO fills possible. Every CRM blank (C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5) is also blank in its ZI row, and C-EE9FFB has no ZI row. Needs a new source; no value invented here.

DISAGREEMENTS (both sides populated):
- Industry, tech family (CRM vs ZI "Computer Software"): C-66D1FC "tech", C-44EA29 "tech", C-60C75F "tech", C-425E2A "Tech ", C-EC3025/C-92D97D/C-D04904/C-77A95A/C-AA8DDA/C-B25F40 "Technology".
  Recommendation: adopt the enrichment (ZoomInfo) taxonomy as canonical for industry — it is the standardized vendor taxonomy; map "tech"/"Tech "/"Technology" → "Computer Software" and "health care" → "Healthcare" (C-7BBDFA, C-50D386 agree with ZI but are non-standard in both). C-0A092933 "SaaS" vs C-0A092934 "Technology" resolves to the survivor's value only after the globex.io merge — flag for the same mapping.
- Country format (US/USA vs ZI "United States"): C-66D1FC, C-950043, C-EC3025, C-96039F, C-77A95A, C-B23205, C-E51FB7, C-D0662E, C-425E2A, C-2D7423. No substantive country conflicts — formatting only. Recommendation: standardize to one code (ISO-2, e.g. "US"/"GB"); either source is fine since values agree in meaning.
- Employee counts: CRM and ZI agree on every matched row (checked all 17 overlapping values) — no conflicts. The only employee conflict is internal: acme-corp.com 500 vs 510 (no ZI row to arbitrate).

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
6. TOP 10 FIXES BY PIPELINE AT STAKE
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Caveat repeated: no deal amounts exist in the extract; ranked by proxy = breadth of affected company/contact records.

1. Obtain and audit deals.csv — 100% of pipeline is currently unauditable on owner/stage/amount/close date/why-buys. Every other ranking is provisional until this lands.
2. Fill employee_count for 8 companies from ZI (=400 each): C-EC3025, C-96039F, C-44EA29, C-D04904, C-B23205, C-60C75F, C-7BBDFA, C-50D386 (16 contacts across them). Source-backed, zero-invention, largest clean win.
3. Backfill contact title/persona — 22 contacts (42.3%) missing buying-role signals. No enrichment source for contacts was provided, so this needs a research/enrichment pass, not a file fill.
4. hq_country gaps — 6 companies (C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB), zero fillable from ZI. Blocks geo/territory segmentation; requires new sourcing.
5. Merge acme-corp.com cluster — survivor C-0A092931, absorb C-0A092932; unresolved 500-vs-510 employee conflict. Duplicate records risk split pipeline if deals attach to either record.
6. Merge globex.io cluster — survivor C-0A092933, absorb C-0A092934; resolve SaaS-vs-Technology via the industry mapping.
7. Repair 4 malformed emails (CT-0010, CT-0080, CT-0081, CT-0192) using the contacts' own domain column as the candidate, with deliverability verification — 4 contacts currently unreachable.
8. Close the enrichment coverage gap — 7 domains with no ZI row (acme-corp.com, globex.io, ba969b.com, 332637.com, 93c8bf.com, ee9ffb.com, c9bb20.com), covering 9 CRM records incl. duplicates. C-93C8BF's employee_count and all six hq_country gaps partially live here.
9. Standardize industry taxonomy — 10 tech-family records + 2 "health care" records + trailing-space "Tech " values (C-425E2A, C-BA969B, C-93C8BF, C-C9BB20). Blocks ICP/segment filters.
10. Resolve CT-0011 domain mismatch (user1@other-domain.com at C-66D1FC) — verify legitimacy or correct; a mis-routed contact on an active buying group.

Honorable mention: 14 companies (41.2%) have zero associated contacts — flag for coverage review once deals confirm which of them carry open pipeline.
deal-intelligence 4 tests
closed-lost-classification0.60
Plain-English question: Why did we lose the last six months of deals — one primary category per deal, and whose side is the reason on?
Prompt — what the model saw
Classify every deal lost in the last 6 months below into one primary category: pricing, competitor, no decision, timing, product gap, champion left, or other. Also assign a side (Bonusly, buyer, unknown). Then summarize: category counts, the side split, the number of deals where the structured closed_lost_tag clearly disagrees with the free-text reason, and the two patterns most worth acting on. Base everything only on the tag and free-text provided.
cost $0.0379 · 399s · in 573 / out 23,715 tok
✗ total_90✗ competitor_top✓ category_counts_present✓ side_split✓ two_patterns

✓ Correct Answer (ground truth)

Plain-English answer: Every closed-lost deal gets exactly one category (pricing, competitor, no decision, timing, product gap, champion left, other) plus a side, per the computed classification.
The answer in one line: total_lost_6mo = 90; top_reason = competitor (26)
total_lost_6mo90
categoriespricing, competitor, no decision, timing, product gap, champion left, other
patternstwo most actionable, tag-vs-text disagreements counted
sideBonusly / buyer / unknown per deal
top_reasoncompetitor (26)
graded checkstotal_90 · competitor_top · category_counts_present · side_split · two_patterns

Model's actual answer

show the model's full answer
METHOD (stated up front)
- Primary category = the main driver in the free-text reason; the structured tag is used as support/tiebreak. When text and tag conflict, the text's stated driver wins the classification and the conflict is counted in the disagreement tally.
- Side rule: Bonusly = text cites a deficiency in Bonusly's product/price/value. Buyer = buyer's internal constraints or choices (budget, timing, priorities, unresponsiveness, incumbent lock-in, leadership decisions). Unknown = text gives no usable driver.
- "No decision" covers: went MIA/unresponsive, deprioritized/not a priority, build-internal, no current R&R need, approval never obtained.

CLASSIFICATION (all 89 deals; alias + side)

TIMING (20) — all buyer
Deal-DB0AAC, Deal-91A056, Deal-29326C, Deal-831B7B, Deal-39E25C, Deal-B3ABED, Deal-B6AC09, Deal-E6E80A, Deal-B038F0, Deal-175756, Deal-BB78F3, Deal-15DA99, Deal-F4AF5D, Deal-79B7A1, Deal-9F176A, Deal-5E64CE (locked into Nectar contract to Oct 2027, plans to revisit), Deal-69CF3D, Deal-ECBF89, Deal-D1A623, Deal-55867E

NO DECISION (32) — all buyer
Deal-AC944F, Deal-214060, Deal-13E9CF (deprioritized; explicitly not budget), Deal-21B045, Deal-ED9AE7 (timing+budget+authority), Deal-988493, Deal-F308CA, Deal-4664E1, Deal-E74A73 (testing points manually first), Deal-D48E0B, Deal-583ADB, Deal-8E27DA (went with swag provider only, didn't want R&R), Deal-E0441F, Deal-7CB44D, Deal-FAC17C (no final approval from Exec IT Director), Deal-50E5D8, Deal-AFA56C, Deal-413C56, Deal-2A292B (building internally), Deal-D1AABF, Deal-FEDBCB, Deal-2BBA21, Deal-7FBAC6, Deal-2FEDDB, Deal-3F86A0, Deal-096750, Deal-F325A5 (layoffs + leadership change), Deal-79E61A, Deal-AE7C4E, Deal-DAB4F1, Deal-B4B50F, Deal-5885B9

COMPETITOR (25)
- Buyer-side, incumbent/lock-in/preference (10): Deal-422BA6 (ADP TotalSource preferred partner), Deal-1BCA50 (budget + stakeholder already down path with another vendor), Deal-A2C349 (staying with Awardco), Deal-8A0992 (Canadian provider), Deal-D0C698 (past Kudos user), Deal-47F1A1 (staying with WorkTango 12 months), Deal-BF2A98 (HiThrive already deployed), Deal-369281 (Paylocity), Deal-9FCD0D (Canadian company per CEO), Deal-64B19A (likely Motivosity)
- Bonusly-side, cited product/value deficiency (4): Deal-F97C37 (other vendor more diversified), Deal-DDAB52 (Rippl: more at same cost, no FX friction), Deal-242273 (digitizing internal points currency was the differentiator), Deal-DC77FE (more customization, e.g. points labeled as dollars; price explicitly not a factor)
- Unknown, no driver given (11): Deal-F7F635, Deal-381C8C, Deal-F1E8A6, Deal-ACE061, Deal-2D2F8D, Deal-0F96AA, Deal-7CC678, Deal-C7156E, Deal-EECC02, Deal-1E7DA9, Deal-286F9C

PRICING (5) — all buyer
Deal-7ED004 (no budget approval), Deal-7B2236 (budget + wants simpler/cheaper), Deal-C33D91 (budget cuts), Deal-DAFB82 (budget needed elsewhere until 2028), Deal-8A119B (no approval)

PRODUCT GAP (4) — all Bonusly
Deal-9048EB (bad fit, multiple feature gaps), Deal-3618CC (wanted Surveys), Deal-5AD03E (wanted more defined budget access), Deal-981AD4 (UI fit + not UK focused)

OTHER (3)
Deal-5DB9B0 (unknown — "Spam.", not ICP), Deal-70F704 (buyer — narrow use case: anniversary awards only, then MIA), Deal-ABD14C (unknown — flat "not interested", unexplained)

CHAMPION LEFT (0) — no deal in this file cites a champion/DM departure.

SUMMARY

Category counts (n=89):
- No decision: 32 (36.0%)
- Competitor: 25 (28.1%)
- Timing: 20 (22.5%)
- Pricing: 5 (5.6%)
- Product gap: 4 (4.5%)
- Other: 3 (3.4%)
- Champion left: 0
Check: 32+25+20+5+4+3+0 = 89 ✓

Side split:
- Buyer: 68 (76.4%) = 20 timing + 32 no decision + 10 competitor + 5 pricing + 1 other
- Unknown: 13 (14.6%) = 11 competitor + 2 other
- Bonusly: 8 (9.0%) = 4 competitor + 4 product gap
Check: 68+13+8 = 89 ✓

Tag vs. free-text disagreements: 4
1. Deal-13E9CF — tag "Doing nothing/Not a priority/Cost"; text says "Not a budget issue" (contradicts the cost component; driver is deprioritization).
2. Deal-9048EB — tag "MIA"; text cites "bad fit based on their desired setup and multiple feature gaps" (tag attributes to silence what the rep attributes to fit).
3. Deal-3618CC — tag "Lost DM"; text says "Wanted Surveys" (feature want, no decision-maker issue stated).
4. Deal-8E27DA — tag "Feature Request"; text says they "didn't want R&R, currently" and went with a swag provider (negates any feature driver).
Caveat (not counted as disagreement since not contradicted): several Competitor tags have no competitor evidence in text — Deal-7CC678 ("Nothing specific provided."), Deal-381C8C, Deal-F1E8A6, and generic "went another direction" texts (Deal-F7F635, Deal-2D2F8D, Deal-EECC02, Deal-C7156E, Deal-1E7DA9). Deal-55867E's Timing tag is likewise unsupported beyond "at this time."

TWO PATTERNS MOST WORTH ACTING ON

1. The "revisit-next-year" timing cohort is recoverable revenue with dates attached (20 deals, 22.5%; $263,211 by my sum: 5115+2975+6300+7200+3360+40001+3000+24000+2340+2880+6600+19600+5760+25000+54600+3360+11520+7200+25200+7200 = 263,211). Nearly every text names a concrete return window (early 2027, Q2 2027, end of year, 2028) and several explicitly want Bonusly later (Deal-DAFB82 "loves Bonusly and is going to loop back"; Deal-BB78F3 "still something we're interested in... as an end goal"). Action: a structured recycle cadence with scheduled re-engagement dates per deal, rather than ad-hoc follow-up.

2. The never-engaged block is a process problem, not a loss-reason problem (32 no-decision deals, 36% — the largest bucket; most tagged MIA with texts like "no contact since intro", "ignored four nudges", "stale when I inherited it"). Deals are reaching formal closed-lost without a real buying conversation. Action: tighten qualification so intro-only deals can't age silently to closed-lost, require multi-threaded/DM contact before later stages, and audit inherited-stale deals at handoff (e.g., Deal-E0441F).

Runner-up not counted in the two: buyer-side incumbent lock-ins (10 competitor deals — Awardco, WorkTango, Kudos, Motivosity, Paylocity, HiThrive, Nectar, ADP-preferred partner, Canadian-preference x2) suggest a displacement/switching-cost playbook (the Nectar exit-fee case, Deal-5E64CE, shows contract-end timing as the entry wedge).
call-transcript-extractionPASS 1.00
Plain-English question: Extract the CRM write-back fields from each call transcript as JSON.
Prompt — what the model saw
For each transcript, extract CRM write-back fields as JSON: why-buys (prospect statements only), pain points, stakeholders from the speaker list, budget signal (prospect-stated or null), timeline signal, competitor mentioned (only if the prospect raised it), next step (only if explicitly agreed), objections, confidence. Never fill a field from what the rep said or hoped.
cost $0.0162 · 155s · in 847 / out 8,598 tok
✓ json_emitted✓ budget_signal✓ null_discipline✓ stakeholders✓ competitor_prospect_only

✓ Correct Answer (ground truth)

Plain-English answer: Per transcript: prospect-only why-buys, pain points, stakeholders, budget and timeline signals, and competitors — e.g. TX-001: '$40k earmarked', live before open enrollment.
The answer in one line: rep = Alex Franklin
allowed_newTX-001, TX-002, TX-003, TX-004, TX-005, TX-006, Achievers, Bucketlist, Nectar
per_transcript{"TX-001": {"why_buys": ["automating anniversary and birthday awards"], "pain_points": ["HR team of three cannot keep up with it manually", "track everything in a spreadsheet, and people slip through
repAlex Franklin
rep_sourced_nulls{"TX-003": ["budget_signal \u2014 only the rep mentioned pricing/budget"], "TX-004": ["next_step \u2014 rep proposed, prospect said 'no promises'"], "TX-006": ["budget_signal \u2014 rep decoy, prospec
graded checksjson_emitted · budget_signal · null_discipline · stakeholders · competitor_prospect_only

Model's actual answer

show the model's full answer
```json
[
  {
    "transcript_id": "TX-001",
    "deal_alias": "Deal-CFE7F4",
    "why_buys": [
      "Automate anniversary and birthday awards — HR team of three cannot keep up with it manually (Prospect (VP People))"
    ],
    "pain_points": [
      "Everything tracked in a spreadsheet; people slip through the cracks (Prospect (HR Admin))"
    ],
    "stakeholders": [
      "Prospect (VP People)",
      "Prospect (HR Admin)"
    ],
    "budget_signal": "$40k earmarked for engagement tools this fiscal year (Prospect (VP People))",
    "timeline_signal": "Wants it live before open enrollment in November (Prospect (VP People))",
    "competitor_mentioned": "Achievers — looked at last year, 'too heavy for a team our size' (raised by Prospect (VP People))",
    "next_step": "Security review with IT lead on September 12 — explicitly agreed (Prospect (VP People))",
    "objections": [
      "SSO and audit logs required for IT to sign off (Prospect (HR Admin))"
    ],
    "confidence": {
      "level": "High",
      "basis": "Prospect-stated budget, timeline, and agreed next step; prior competitor evaluated and rejected."
    }
  },
  {
    "transcript_id": "TX-002",
    "deal_alias": "Deal-70BB30",
    "why_buys": [
      "Tie recognition to retention for the hourly workforce — regretted turnover there is over 30% (Prospect (Head of Total Rewards))"
    ],
    "pain_points": [
      "Regretted turnover in hourly workforce over 30% (Prospect (Head of Total Rewards))"
    ],
    "stakeholders": [
      "Prospect (Head of Total Rewards)",
      "Prospect (CFO)"
    ],
    "budget_signal": "$25k pilot budget approved by finance for this quarter (Prospect (CFO))",
    "timeline_signal": "Decision by end of September (Prospect (CFO))",
    "competitor_mentioned": null,
    "next_step": "Send pilot agreement; prospect will route it to legal this week — explicitly agreed (Prospect (CFO))",
    "objections": [
      "Workday integration 'has to be rock solid' — CFO's stated one condition (Prospect (CFO))"
    ],
    "confidence": {
      "level": "High",
      "basis": "Approved pilot budget, decision deadline, agreed next step; prospect states no other vendor demos yet."
    }
  },
  {
    "transcript_id": "TX-003",
    "deal_alias": "Deal-530B50",
    "why_buys": [
      "Make recognition visible across the 12 retail locations (Prospect (People Ops Manager))"
    ],
    "pain_points": [
      "Recognition not visible across 12 retail locations (Prospect (People Ops Manager))",
      "Store managers have zero budget autonomy for on-the-spot recognition today (Prospect (People Ops Manager))"
    ],
    "stakeholders": [
      "Prospect (People Ops Manager)"
    ],
    "budget_signal": null,
    "timeline_signal": "No rush on prospect's side until Q1 (Prospect (People Ops Manager))",
    "competitor_mentioned": "Bucketlist — CEO used it at her last company and liked it (raised by Prospect (People Ops Manager))",
    "next_step": "Schedule a call with the CEO; prospect will send two times — explicitly agreed (Prospect (People Ops Manager))",
    "objections": [
      "CEO has to be sold first — she decides anything people-related (Prospect (People Ops Manager))"
    ],
    "confidence": {
      "level": "Low-Medium",
      "basis": "No prospect-stated budget (the $8/employee/month figure was rep-stated only), no urgency until Q1, decision gate at a CEO not yet engaged, and a favorable prior competitor experience."
    }
  },
  {
    "transcript_id": "TX-004",
    "deal_alias": "Deal-180D02",
    "why_buys": [
      "Consolidate three separate recognition tools into one (Prospect (VP People))"
    ],
    "pain_points": [
      "Paying for three tools and none of them talk to the HRIS (Prospect (VP People))",
      "Procurement cycle runs six to eight weeks minimum (Prospect (IT Security Lead))",
      "Prior vendor's security review took three months (Prospect (IT Security Lead))"
    ],
    "stakeholders": [
      "Prospect (VP People)",
      "Prospect (IT Security Lead)"
    ],
    "budget_signal": "Can approve without going to the board if under $15k annually (Prospect (VP People)) — approval threshold stated, not a budget amount",
    "timeline_signal": "Procurement cycle six to eight weeks minimum (Prospect (IT Security Lead))",
    "competitor_mentioned": null,
    "next_step": null,
    "objections": [
      "Security review took three months for the last vendor — stated hesitation (Prospect (IT Security Lead))",
      "CFO follow-up proposed by rep not committed — 'Maybe — I need to check her calendar, no promises' (Prospect (VP People))"
    ],
    "confidence": {
      "level": "Medium",
      "basis": "Clear consolidation motive and stated self-approval threshold, but no explicitly agreed next step and a long procurement/security-review cycle."
    }
  },
  {
    "transcript_id": "TX-005",
    "deal_alias": "Deal-F8767A",
    "why_buys": [
      "Automate service milestones (Prospect (HR Director))",
      "Analytics on recognition equity across departments (Prospect (HR Director))"
    ],
    "pain_points": [
      "Night-shift teams feel invisible — their engagement scores run 20 points lower (Prospect (People Ops Coordinator))",
      "Exec team is skeptical after a failed rollout two years ago (Prospect (HR Director))",
      "Mid-pilot with Nectar — incumbent experience to beat (Prospect (HR Director))"
    ],
    "stakeholders": [
      "Prospect (HR Director)",
      "Prospect (People Ops Coordinator)"
    ],
    "budget_signal": "$12k approved under the engagement line (Prospect (HR Director))",
    "timeline_signal": "Needs to be running before the January all-hands (Prospect (HR Director))",
    "competitor_mentioned": "Nectar — prospect is mid-pilot with them (raised by Prospect (HR Director))",
    "next_step": "Present to the exec team on October 2 — explicitly agreed (Prospect (HR Director))",
    "objections": [
      "Exec team skepticism after a failed rollout two years ago (Prospect (HR Director))",
      "Must beat the Nectar pilot experience (Prospect (HR Director))"
    ],
    "confidence": {
      "level": "Medium-High",
      "basis": "Approved budget, deadline, and agreed exec presentation; offset by an active Nectar pilot and exec skepticism."
    }
  },
  {
    "transcript_id": "TX-006",
    "deal_alias": "Deal-EE195F",
    "why_buys": [
      "Cut the admin time on service awards (Prospect (HR Manager))"
    ],
    "pain_points": [
      "Personally spends five hours a month ordering and shipping plaques (Prospect (HR Manager))",
      "COO usually prefers building things in-house (Prospect (HR Manager))"
    ],
    "stakeholders": [
      "Prospect (HR Manager)"
    ],
    "budget_signal": "Prospect states budget is not a constraint — 'Budget isn't the issue — time is' (Prospect (HR Manager)); no amount stated",
    "timeline_signal": "Fiscal year resets in February, so a Q1 start is realistic (Prospect (HR Manager))",
    "competitor_mentioned": "None named — prospect states 'Nobody else'; the only stated comparison is building internally, which is not a vendor (Prospect (HR Manager))",
    "next_step": "Send the one-page overview; prospect will forward it to the COO this week — explicitly agreed (Prospect (HR Manager))",
    "objections": [
      "COO usually prefers building things in-house — internal build is the stated alternative (Prospect (HR Manager))"
    ],
    "confidence": {
      "level": "Medium",
      "basis": "No stated budget constraint, Q1 timeline, and agreed next step; but single-threaded through the HR Manager with a build-biased COO as the gate."
    }
  }
]
```
next-to-close0.75
Plain-English question: Which 3 deals are closest to signature right now?
Prompt — what the model saw
Which 3 deals are closest to signature right now? Use the CRM data and the Slack thread; Slack often has the real status before the CRM does. For each: deal, amount, why it is close, what is left. Under 250 words.
cost $0.0087 · 69s · in 182 / out 3,532 tok
✗ exactly_three✓ top3_named✓ slack_first✓ blocked_excluded

✓ Correct Answer (ground truth)

Plain-English answer: Deal-547B2B ($11,200), Deal-B7EBD1 ($9,000), and Deal-403845 ($9,000) — all DS5 COMMIT deals closing September 10-11, with what is left on each.
blocked_excludeddeal_alias=Deal-2465CE; reason=champion left + procurement freeze (Slack #sales-sparring 2026-09-04)
exactly_three1
ranking_ruleSlack signature-imminent deal ranks first; remaining ranked by stage (DS5>DS4), forecast (COMMIT>BEST_CASE>PIPELINE), close_date asc, amount desc
top3rank=1; deal_alias=Deal-547B2B; amount=11200; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11 | rank=2; deal_alias=Deal-B7EBD1; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-10 | rank=3; deal_alias=Deal-403845; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11
graded checksexactly_three · top3_named · slack_first · blocked_excluded

Model's actual answer

show the model's full answer
Closest to signature (Slack thread dated 2026-09-04 is the most recent signal; CRM close dates are 09-10/09-11):

1. Deal-547B2B — $11,200 (Alex Franklin, DS5/COMMIT, CRM close 2026-09-11)
   Why close: Slack — redlines came back clean, signing page is out, their VP People said they are signing tomorrow; "signature-imminent."
   What's left: countersignature only.

2. Deal-403845 — $9,000 (Alex Franklin, DS5/COMMIT, CRM close 2026-09-11)
   Why close: Slack — order form is with their finance team; Dana confirms it's "moving fine."
   What's left: finance sign-off → signature.

3. Deal-A2B47C — $6,360 (Alex Franklin, DS5/COMMIT, CRM close 2026-09-11)
   Why close: Slack — "still warm, just normal legal-review pace."
   What's left: legal review completion → signature.

Key Slack-over-CRM correction: Deal-2465CE ($5,400, DS5/COMMIT in CRM, close 09-10) should NOT be counted — Slack says the champion left, procurement froze new vendors, and Dana is pulling it from commit; realistically a Q4 deal now.

Also noted: Deal-D348E1 ($13,770) is warm per Slack but closes 10-15, a month out — not top-3. Deal-B7EBD1 ($9,000, close 09-10) has no Slack signal, so its actual status can't be verified beyond the CRM row.
product-gap-detectionPASS 1.00
Plain-English question: Which prospect complaints are true product gaps versus rollout or enablement issues?
Prompt — what the model saw
Surface true product gaps raised by prospects in active deals. For each candidate: quote the prospect line, classify as real gap, plan gate, or rollout/enablement issue using the product docs excerpt, and name the deal and amount. Only real gaps go in the summary. Rep language does not count as prospect voice.
cost $0.0074 · 47s · in 0 / out 2,707 tok
✓ real_gaps✓ plan_gate_classified✓ enablement_classified✓ gap_summary

✓ Correct Answer (ground truth)

Plain-English answer: Two real gaps (SCIM provisioning, ADP Workforce Now integration) separated from plan gates and enablement issues, each with the prospect quote and deal.
allowed_newTG-001, TG-002, TG-003, TG-004, TG-005, SCIM user provisioning, ADP Workforce Now integration, Custom report builder
enablement_issuestranscript_id=TG-004; deal_alias=Deal-84DBA6; amount=16000; feature=Slack integration recognition adoption; quote=We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.; why_enablement=Slack integration exists on all plans; managers were never trained
excluded_rep_voicetranscript_id=TG-005; line=Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap.; reason=rep voice, not prospect-raised
plan_gatestranscript_id=TG-003; deal_alias=Deal-CFE7F4; amount=32000; feature=Custom report builder; tier_required=Enterprise; quote=I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?
real_gapstranscript_id=TG-001; deal_alias=Deal-EC3025; amount=62000; feature=SCIM user provisioning; quote=We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.; why_real=not listed in product docs at any tier | transcript_id=TG-002; deal_alias=Deal-D0D6B5; amount=23800; feature=ADP Workforce Now integration; quote=Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.; why_real=supported HRIS list is Workday/BambooHR/Gusto/Rippling only
graded checksreal_gaps · plan_gate_classified · enablement_classified · gap_summary

Model's actual answer

show the model's full answer
CANDIDATE REVIEW (per transcript)

1) Deal-EC3025 (TG-001) — Prospect (IT Security Lead)
Quote: "We need SCIM user provisioning — our security team will not approve any tool that requires manual account management."
Classification: REAL GAP
Basis: Product docs state SCIM user provisioning is "NOT currently listed as supported capabilities" — absent at every tier, so not a plan gate.
Amount: NOT PROVIDED in source data.

2) Deal-D0D6B5 (TG-002) — Prospect (HRIS Manager)
Quote: "Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us."
Classification: REAL GAP
Basis: Docs state ADP Workforce Now is "NOT currently listed as supported capabilities"; supported HRIS list is Workday, BambooHR, Gusto, Rippling only.
Amount: NOT PROVIDED in source data.

3) Deal-CFE7F4 (TG-003) — Prospect (VP People)
Quote: "I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?"
Classification: PLAN GATE
Basis: Docs list "Custom report builder" as Enterprise-only. Capability exists; blocked by tier, not absent from product.
Amount: NOT PROVIDED in source data.

4) Deal-84DBA6 (TG-004) — Prospect (People Ops Manager)
Quote: "We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it."
Classification: ROLLOUT/ENABLEMENT ISSUE
Basis: Docs show Slack integration supported on all plans (Core/Pro/Enterprise: yes). Feature works; this is adoption/training, not a product gap.
Amount: NOT PROVIDED in source data.

5) Deal-36C33F (TG-005) — EXCLUDED
The only gap statement ("no native mobile app for hourly workers yet") is rep language (Alex Franklin), not prospect voice. Per instructions, rep language does not count. The prospect's sole line accepts the workaround: "Good to know. The web version should be fine for our office staff for now." No prospect-raised gap.

COUNT CHECK
- Transcripts: 5
- Prospect-raised candidates: 4 (TG-001, TG-002, TG-003, TG-004); TG-005 excluded (rep-raised): 4 + 1 = 5
- Classifications: 2 real + 1 plan gate + 1 enablement = 4 ✓

SUMMARY — TRUE PRODUCT GAPS (prospect-raised, active deals)

1. Deal-EC3025 — SCIM user provisioning.
   Quote: "We need SCIM user provisioning — our security team will not approve any tool that requires manual account management."
   Docs confirm SCIM is not a supported capability at any tier.

2. Deal-D0D6B5 — ADP Workforce Now integration.
   Quote: "Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us."
   Docs confirm ADP Workforce Now is not a supported integration at any tier.

MISSING DATA (explicit)
- Deal amounts: the source CSV has no amount field for any deal (Deal-EC3025, Deal-D0D6B5, Deal-CFE7F4, Deal-84DBA6, Deal-36C33F). Amounts cannot be stated without inventing numbers.
- "Active deal" status is not stated in the source data; taken as given by the task framing.
rep-performance 4 tests
stale-pipeline-by-repPASS 1.00
Plain-English question: Which open deals have had no email, call, or meeting in the last 7 days?
Prompt — what the model saw
The snapshot date is 2026-09-05. Using the open deals and the engagements-by-deal table, list every open deal with no logged email, call, or meeting in the last 7 days. Compute recency from the engagements table's last_email/last_call/last_meeting fields, NOT from any date field on the deal (the deal's last_contacted_field is unreliable). For each stale deal give: deal alias, owner name, stage, amount, days since last contact. Group by owner, order by amount descending within each owner. Then state, per owner, how many stale deals they carry and the total stale amount. (Do not write outreach text.)
cost $0.0506 · 431s · in 17,417 / out 24,838 tok
✓ stale_recall_ge_80✓ grouped_by_owner✓ jake_top_amount✓ kevin_stale_amt✓ jenn_stale_amt✓ stale_count

✓ Correct Answer (ground truth)

Plain-English answer: 65 open deals are stale; the largest is Deal-2D1F1B at $240,000, with Bryce Harmon carrying the most stale dollars at $626,243.
The answer in one line: jake_amt = 626,243; jenn_amt = 261,645; kevin_amt = 252,905.03; top_amount = 240,000; stale_count = 65; top_alias = Deal-2D1F1B
jake_amt626,243
jenn_amt261,645
kevin_amt252,905.03
stale_count65
top_amount240,000
top_aliasDeal-2D1F1B
graded checksstale_recall_ge_80 · grouped_by_owner · jake_top_amount · kevin_stale_amt · jenn_stale_amt · stale_count

Model's actual answer

show the model's full answer
Method (stated up front):
- Staleness window: last 7 days = 2026-08-29 through 2026-09-05. A deal is stale if none of last_email / last_call / last_meeting falls in that window.
- Days since last contact = 2026-09-05 minus the most recent contact date on or before the snapshot (max of the three engagement fields, excluding future-dated values). Example: 2026-09-05 − 2026-08-20 = 16 days.
- Deals with no row in engagements_by_deal_90d.csv have no logged email/call/meeting at all → stale, days not computable (marked n/a).
- Caveat: 10 deals have a future-dated last_meeting (2026-09-09 to 2026-10-02, after the snapshot — likely scheduled meetings). I treated those as not-yet-occurred and used the most recent past contact. They are marked *. If you'd rather count future-dated meetings as contact, exactly these 10 drop out: Deal-944310, Deal-7599B8, Deal-3795AD, Deal-A414F6, Deal-C1FA6D, Deal-927338, Deal-01E193, Deal-819506, Deal-8BA24E, Deal-8FDCD2.

No deal had last contact exactly on 2026-08-29 or 08-30, so the window boundary doesn't affect results.

Owners ordered by total stale amount desc.

=== BRYCE HARMON — 18 stale deals, $692,964 ===
Deal-2D1F1B | DS1 | 240000   | last mtg 2026-06-16 | 81d
Deal-66D1FC | DS1 | 99000    | last email 2026-08-20 | 16d
Deal-950043 | DS1 | 70000    | last email 2026-08-17 | 19d
Deal-B23205 | DS1 | 45000    | last email 2026-08-20 | 16d
Deal-7BBDFA | DS3 | 37440    | last email 2026-07-21 | 46d
Deal-332637 | DS2 | 36000    | last email 2026-08-27 | 9d
Deal-1BEEBF | DS1 | 31500    | last email 2026-08-17 | 19d
Deal-A414F6*| DS1 | 25200    | last email 2026-08-17 (mtg 09-10 future) | 19d
Deal-C5658B | DS1 | 23400    | last email 2026-08-20 | 16d
Deal-40522D | DS3 | 21000    | last email 2026-08-17 | 19d
Deal-C1FA6D*| DS1 | 18000    | last email 2026-08-20 (mtg 09-15 future) | 16d
Deal-01E193*| DS1 | 12600    | last email 2026-08-28 (mtg 09-09 future) | 8d
Deal-F0EBBB | DS3 | 11400    | last email 2026-08-12 | 24d
Deal-927338*| DS1 | 10920    | last email 2026-08-18 (mtg 09-17 future) | 18d
Deal-E25A09 | DS1 | 6000     | last email 2026-08-27 | 9d
Deal-C9C286 | DS2 | 5502     | last email 2026-08-27 | 9d
Deal-012CB1 | DS1 | 1        | last email 2026-08-13 | 23d
Deal-3795AD*| DS2 | 1        | last email 2026-08-28 (mtg 10-02 future) | 8d
Sum check: 240000+99000+70000+45000+37440+36000+31500+25200+23400+21000+18000+12600+11400+10920+6000+5502+1+1 = 692,964

=== DANA MERCER — 16 stale deals, $279,495 ===
Deal-44EA29 | DS2 | 60000    | last email 2026-08-26 | 10d
Deal-E51FB7 | DS2 | 43875    | last email 2026-08-18 | 18d
Deal-B42F46 | DS1 | 27000    | last email 2026-08-17 | 19d
Deal-BA3DDC | DS3 | 23400    | last email 2026-08-20 | 16d
Deal-9DDE86 | DS2 | 20000    | last email 2026-08-21 | 15d
Deal-215CCA | DS3 | 18900    | last mtg 2026-08-19 | 17d
Deal-5EED42 | DS3 | 16250    | last email/call 2026-08-25 | 11d
Deal-57887A | DS2 | 15000    | last email 2026-08-28 | 8d
Deal-944310*| DS4 | 10500    | last email 2026-08-03 (mtg 09-15 future) | 33d
Deal-B7EBD1 | DS5 | 9000     | last email 2026-08-20 | 16d
Deal-3974EB | DS4 | 9000     | last email 2026-08-28 | 8d
Deal-F40F04 | DS2 | 8100     | last email 2026-08-21 | 15d
Deal-7599B8*| DS3 | 7350     | last email 2026-08-18 (mtg 09-10 future) | 18d
Deal-87DDD1 | DS1 | 5000     | last email 2026-08-17 | 19d
Deal-F336B6 | DS3 | 4200     | last email 2026-08-21 | 15d
Deal-0660B4 | DS4 | 1920     | last email 2026-08-10 | 26d
Sum check: 60000+43875+27000+23400+20000+18900+16250+15000+10500+9000+9000+8100+7350+5000+4200+1920 = 279,495

=== COLE INGRAM — 18 stale deals, $252,905.03 ===
Deal-D04904 | DS2 | 58529.25 | last email 2026-08-25 | 11d
Deal-B25F40 | DS3 | 40000    | last email 2026-08-28 | 8d
Deal-813836 | DS2 | 32175    | last email 2026-08-25 | 11d
Deal-1BA595 | DS2 | 31750    | last email 2026-08-25 | 11d
Deal-CFE1E8 | DS3 | 18000    | last email 2026-08-25 | 11d
Deal-CD47A6 | DS2 | 12168    | last email 2026-08-25 | 11d
Deal-627646 | DS3 | 11193    | last email 2026-08-25 | 11d
Deal-FF809F | DS2 | 7781.20  | last email 2026-08-25 | 11d
Deal-AF932D | DS2 | 7225.40  | last email 2026-08-25 | 11d
Deal-A71728 | DS2 | 6947.50  | last email 2026-08-25 | 11d
Deal-8BC9F5 | DS2 | 5616     | last email 2026-08-26 | 10d
Deal-175395 | DS3 | 4779.88  | last email 2026-08-25 | 11d
Deal-481E24 | DS3 | 4140     | last call 2026-08-26 | 10d
Deal-C7F9BF | DS2 | 3360     | last email 2026-08-25 | 11d
Deal-2F3A66 | DS3 | 3334.80  | last email 2026-08-25 | 11d
Deal-342E96 | DS2 | 2700     | last email 2026-08-12 | 24d
Deal-E568D5 | DS3 | 1875     | last email 2026-08-25 | 11d
Deal-FD9F4E | DS5 | 1330     | last email 2026-08-26 | 10d
Sum check: 58529.25+40000+32175+31750+18000+12168+11193+7781.20+7225.40+6947.50+5616+4779.88+4140+3360+3334.80+2700+1875+1330 = 252,905.03

=== ALEX FRANKLIN — 20 stale deals, $113,936 ===
Deal-CC08D1 | DS1 | 24000    | last email 2026-08-20 | 16d
Deal-E73427 | DS3 | 18000    | last email 2026-08-26 | 10d
Deal-885F45 | DS2 | 9300     | last email 2026-08-24 | 12d
Deal-C2FF3C | DS1 | 8316     | last email 2026-08-26 | 10d
Deal-3EED2C | DS2 | 7200     | no row in engagements table | n/a
Deal-0D2F7A | DS3 | 5100     | last call 2026-08-24 | 12d
Deal-6C60D4 | DS3 | 4800     | last call 2026-08-24 | 12d
Deal-13FEBD | DS2 | 4680     | last call 2026-08-24 | 12d
Deal-819506*| DS1 | 4400     | last email 2026-08-28 (mtg 09-09 future) | 8d
Deal-9D0060 | DS3 | 3840     | last email 2026-08-24 | 12d
Deal-690476 | DS2 | 3600     | last call 2026-08-18 | 18d
Deal-C6D97A | DS4 | 3240     | last email 2026-08-28 | 8d
Deal-EE195F | DS3 | 3120     | last email 2026-08-28 | 8d
Deal-278DEC | DS3 | 2700     | last email 2026-08-28 | 8d
Deal-635B8E | DS3 | 2600     | last email 2026-08-18 | 18d
Deal-6883F3 | DS1 | 2400     | last email 2026-08-20 | 16d
Deal-4A13AD | DS3 | 2160     | last email 2026-08-10 | 26d
Deal-F67D31 | DS2 | 1800     | last email 2026-08-28 | 8d
Deal-5FDCE4 | DS3 | 1600     | last email 2026-08-24 | 12d
Deal-BA571A | DS4 | 1080     | last email 2026-08-18 | 18d
Sum check: 24000+18000+9300+8316+7200+5100+4800+4680+4400+3840+3600+3240+3120+2700+2600+2400+2160+1800+1600+1080 = 113,936

=== FARID OSMAN — 2 stale deals, $2,881 ===
Deal-8BA24E*| DS1 | 2880     | last email 2026-08-28 (mtg 09-15 future) | 8d
Deal-8FDCD2*| DS1 | 1        | last email 2026-08-21 (mtg 09-15 future) | 15d
Sum check: 2880+1 = 2,881

=== ELENA SINCLAIR — 1 stale deal, $2,100 ===
Deal-57FF13 | DS1 | 2100     | no row in engagements table | n/a
(Note: this deal's close_date 2026-09-02 predates the snapshot yet it appears in the open-deals file; included as open per the file.)

=== PER-OWNER SUMMARY ===
Owner          Stale deals   Total stale amount
Bryce Harmon        18          $692,964.00
Dana Mercer         16          $279,495.00
Cole Ingram         18          $252,905.03
Alex Franklin       20          $113,936.00
Farid Osman          2          $2,881.00
Elena Sinclair       1          $2,100.00
TOTAL               75        $1,344,281.03

If the 10 future-dated-meeting deals (*) were instead counted as contacted, the totals would be: Bryce 14 deals / $636,442; Dana 14 / $261,645; Alex 19 / $109,536; Farid 0 / $0; Cole and Elena unchanged; overall 65 deals / $1,262,508.03.
activity-mix-vs-outcomePASS 1.00
Plain-English question: Per rep, what did the last 30 days of emails, calls, and meetings look like next to the deals that advanced?
Prompt — what the model saw
The snapshot date is 2026-09-05. For each sales rep (owner), compute the last-30-day totals of emails, calls, and meetings from the engagements table, and the number of deals that entered DS2 in the last 30 days (t_ds2 within 30 days of the snapshot). Report per rep: the activity mix as percentages (emails/calls/meetings share of that rep's total activities), total activities per DS2 entry, and rank the reps by that efficiency ratio (lowest activities per DS2 entry = most efficient). Name the most efficient rep and the highest-volume rep, and note if they differ. Do not attribute cause.
cost $0.0532 · 521s · in 17,414 / out 26,680 tok
✓ alex_ds2_30d✓ justin_volume_high✓ efficient_named✓ mix_present✓ no_causal_claim

✓ Correct Answer (ground truth)

Plain-English answer: Alex Franklin is the standout signal with 18 deals entering DS2 in 30 days (next closest: Bryce Harmon with 4), alongside per-rep activity totals.
The answer in one line: alex_ds2_30d = 18; bryce_ds2_30d = 4; note = legacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
alex_ds2_30d18
bryce_ds2_30d4
notelegacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
graded checksalex_ds2_30d · justin_volume_high · efficient_named · mix_present · no_causal_claim

Model's actual answer

show the model's full answer
Snapshot: 2026-09-05. Last-30-day window = 2026-08-06 through 2026-09-05 (inclusive of the boundary date 2026-08-06).

Activity totals are the sum of emails_30d / calls_30d / meetings_30d across each owner's deals in deals_open.csv, joined to engagements_by_deal_90d.csv by deal_id.

DATA GAPS (stated up front):
- Deal-3EED2C (owner Alex Franklin, t_ds2=2026-09-03) has no row in engagements_by_deal_90d.csv — counted in DS2 entries, contributes 0 activities.
- Deal-57FF13 (Elena Sinclair) has no row in engagements_by_deal_90d.csv; it also never reached DS2 (t_ds2 blank). Elena: 0 activities, 0 DS2 entries → ratio undefined, excluded from ranking.
- Gavin Porter and Hugo Lindqvist are archived and own no deals in the snapshot → no data.
- Boundary note: three deals have t_ds2 = 2026-08-06 exactly (Deal-1CCE5C Bryce; Deal-EE195F, Deal-D9A72E Alex). Counted as in-window. Excluding them would give Alex 384/17 = 22.6 and Bryce 205/3 = 68.3; the most-efficient rep is unchanged either way.

PER-REP RESULTS

Alex Franklin (84342457) — 67 deals
- Emails 307, Calls 36, Meetings 41 → Total 384
- Mix: 307/384 = 79.9% | 36/384 = 9.4% | 41/384 = 10.7%
- DS2 entries (19): Deal-403845, Deal-1FC049, Deal-3EED2C, Deal-7FA0C3, Deal-E531A6, Deal-5296C9, Deal-36C33F, Deal-EE195F, Deal-F436DA, Deal-317E6F, Deal-D1E6C2, Deal-D9A72E, Deal-CA5E44, Deal-4F775F, Deal-898FC5, Deal-D8ABF7, Deal-46988D, Deal-E73427, Deal-92D97D
- Activities per DS2 entry: 384/19 = 20.2

Bryce Harmon (119337721) — 35 deals
- Emails 162, Calls 0, Meetings 43 → Total 205
- Mix: 162/205 = 79.0% | 0/205 = 0.0% | 43/205 = 21.0%
- DS2 entries (4): Deal-25F752, Deal-CA7DC0, Deal-1CCE5C, Deal-D73B89
- Activities per DS2 entry: 205/4 = 51.25

Cole Ingram (83155924) — 22 deals
- Emails 96, Calls 14, Meetings 1 → Total 111
- Mix: 96/111 = 86.5% | 14/111 = 12.6% | 1/111 = 0.9%
- DS2 entries (2): Deal-42326B, Deal-1BA595
- Activities per DS2 entry: 111/2 = 55.5

Farid Osman (716654662) — 7 deals
- Emails 38, Calls 0, Meetings 34 → Total 72
- Mix: 38/72 = 52.8% | 0/72 = 0.0% | 34/72 = 47.2%
- DS2 entries (1): Deal-499BF6
- Activities per DS2 entry: 72/1 = 72.0

Dana Mercer (83155923) — 24 deals
- Emails 84, Calls 18, Meetings 11 → Total 113
- Mix: 84/113 = 74.3% | 18/113 = 15.9% | 11/113 = 9.7%
- DS2 entries (1): Deal-57887A
- Activities per DS2 entry: 113/1 = 113.0

Elena Sinclair (701163055) — 1 deal (Deal-57FF13)
- No engagement data; 0 DS2 entries → mix and ratio undefined; excluded from ranking.

RANKING (lowest activities per DS2 entry = most efficient)
1. Alex Franklin — 20.2
2. Bryce Harmon — 51.25
3. Cole Ingram — 55.5
4. Farid Osman — 72.0
5. Dana Mercer — 113.0

Most efficient rep: Alex Franklin (20.2 activities per DS2 entry).
Highest-volume rep: Alex Franklin (384 total activities).
They are the same rep. (No cause attributed, per instructions.)
multithreading-gapPASS 1.00
Plain-English question: Which open deals are single-threaded or under-threaded?
Prompt — what the model saw
Find every open deal that is single-threaded (fewer than 2 active contacts) or under-threaded (fewer than 3, or all contacts in one persona). Active means engaged in the last 60 days and not marked former. Personas: economic buyer, champion, HR admin, IT security, finance. For each flagged deal: amount, stage, active contact count, personas present, personas missing, the most valuable persona to add given the stage, and any on-file unengaged contact at that company who fits; otherwise say none on file.
cost $0.0141 · 125s · in 893 / out 7,018 tok
✓ flagged_recall✓ persona_language✓ single_thread_term✓ on_file_fit

✓ Correct Answer (ground truth)

Plain-English answer: Deals with fewer than 2 active contacts (single-threaded) or fewer than 3, or all in one persona (under-threaded), each flagged with the missing persona.
The answer in one line: active_window_start = 2026-07-07
active_window_start2026-07-07
allowed_new59 items — e.g. C-077A0E, C-0D15DF, C-1BDFE0, C-2AE3AA
best_persona_ruleDS2/DS3: economic buyer > IT security > finance > HR admin > champion; DS4/DS5: economic buyer > finance > IT security > HR admin > champion; first persona not present among active contacts
flag_ruleflag if active_contacts < 2 (single-threaded) or active_contacts < 3 or all active contacts share one persona (under-threaded); active = engaged within 60 days and not former
flaggeddeal_alias=Deal-36C33F; amount=15000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-50D386; amount=36000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5408B0; amount=14850; stage=DS4; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5BFE3B; amount=18000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-885F45; amount=9300; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=IT security | deal_alias=Deal-92D97D; amount=60000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer … (+5 more)
healthy_not_flaggedDeal-4B0BEB, Deal-84DBA6, Deal-D348E1
persona_universeeconomic buyer, champion, HR admin, IT security, finance
graded checksflagged_recall · persona_language · single_thread_term · on_file_fit

Model's actual answer

show the model's full answer
Analysis date: 2026-09-24. Active = last_engaged_date on/after 2026-07-26 and is_former=false. Arithmetic for the 60-day cutoff: 2026-09-24 − 24d = 2026-08-31; − 31d = 2026-07-31; − 5d = 2026-07-26.

MISSING DATA (applies to every deal; neither file contains these fields):
- Amount: NOT PROVIDED for any deal.
- Stage: NOT PROVIDED for any deal. Therefore "most valuable persona to add given the stage" cannot be conditioned on stage; I apply a stated heuristic instead: economic buyer if absent; else champion if absent; else next missing persona (IT security / finance / HR admin).
- Open/closed status: no close-date or status column; all 14 deals treated as open.

FLAGGED DEALS (11 of 14)

1) Deal-EC3025 (61032318100, C-FDD0C7) — single-threaded
- Amount/stage: NOT PROVIDED
- Active contacts: 1 — CT-047C54 (champion, 2026-09-02). CT-F2C1AE (economic buyer) excluded: is_former=true.
- Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
- Most valuable add: economic buyer.
- On-file unengaged fit: CT-6827DB, Chief People Officer (economic buyer) — C-FDD0C7.

2) Deal-92D97D (59728118877, C-E23238) — single-threaded
- Amount/stage: NOT PROVIDED
- Active contacts: 1 — CT-01F5B4 (HR admin, 2026-08-28). CT-A902AE (champion) not active: last engaged 2026-06-01, before cutoff.
- Personas present: HR admin. Missing: economic buyer, champion, IT security, finance.
- Most valuable add: champion (zero active champions).
- On-file unengaged fit: none on file (C-E23238 not in unengaged_contacts.csv).

3) Deal-50D386 (61055128146, C-EB10E4) — under-threaded (2 active)
- Amount/stage: NOT PROVIDED
- Active contacts: 2 — CT-AA41B2 (champion, 2026-09-01), CT-B9C35B (HR admin, 2026-08-25).
- Personas present: champion, HR admin. Missing: economic buyer, IT security, finance.
- Most valuable add: economic buyer.
- On-file unengaged fit: CT-A1C4B3, Chief People Officer (economic buyer) — C-EB10E4.

4) Deal-D0D6B5 (60081655042, C-32918E) — under-threaded (all contacts one persona)
- Amount/stage: NOT PROVIDED
- Active contacts: 3 — CT-87CED4 (2026-09-02), CT-DE6D7C (2026-08-19), CT-FD70B2 (2026-08-07); all champion.
- Personas present: champion only. Missing: economic buyer, HR admin, IT security, finance.
- Most valuable add: economic buyer.
- On-file unengaged fit: CT-1FA4DB, Chief People Officer (economic buyer) — C-32918E.

5) Deal-5BFE3B (51674270311, C-535D36) — under-threaded (2 active, both champion)
- Amount/stage: NOT PROVIDED
- Active contacts: 2 — CT-57123B (2026-08-31), CT-5CE757 (2026-08-12); all champion.
- Personas present: champion only. Missing: economic buyer, HR admin, IT security, finance.
- Most valuable add: economic buyer.
- On-file unengaged fit: none on file (C-535D36 not in unengaged_contacts.csv).

6) Deal-36C33F (63739413805, C-077A0E) — single-threaded
- Amount/stage: NOT PROVIDED
- Active contacts: 1 — CT-4FE556 (IT security, 2026-08-15). CT-405B45 (champion) and CT-86B22F (economic buyer) excluded: is_former=true.
- Personas present: IT security. Missing: economic buyer, champion, HR admin, finance.
- Most valuable add: champion (champion layer is former; none on file). Note: economic buyer is also missing and CT-1DB73E, Chief People Officer (economic buyer), is on file unengaged at C-077A0E.

7) Deal-885F45 (60686135564, C-5E8EFB) — under-threaded (2 active)
- Amount/stage: NOT PROVIDED
- Active contacts: 2 — CT-51C81E (economic buyer, 2026-08-26), CT-D9A0E8 (champion, 2026-08-11).
- Personas present: economic buyer, champion. Missing: HR admin, IT security, finance.
- Most valuable add: IT security (EB + champion already in place; stage unavailable so heuristic applied).
- On-file unengaged fit: CT-B3F25D, IT Security Lead (IT security) — C-5E8EFB.

8) Deal-FCBE5B (62639586615, C-737030) — single-threaded
- Amount/stage: NOT PROVIDED
- Active contacts: 1 — CT-4A5317 (champion, 2026-08-29).
- Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
- Most valuable add: economic buyer.
- On-file unengaged fit: none on file (C-737030 not in unengaged_contacts.csv).

9) Deal-5408B0 (60182332309, C-2AE3AA) — under-threaded (2 active)
- Amount/stage: NOT PROVIDED
- Active contacts: 2 — CT-D33AE4 (champion, 2026-09-01), CT-8742FD (HR admin, 2026-08-18).
- Personas present: champion, HR admin. Missing: economic buyer, IT security, finance.
- Most valuable add: economic buyer.
- On-file unengaged fit: CT-07FA76, Chief People Officer (economic buyer) — C-2AE3AA.

10) Deal-C6D97A (62121783047, C-5A8FC2) — under-threaded (all contacts one persona)
- Amount/stage: NOT PROVIDED
- Active contacts: 3 — CT-223DDC (2026-08-31), CT-B03555 (2026-08-20), CT-4E8A2B (2026-08-05); all champion.
- Personas present: champion only. Missing: economic buyer, HR admin, IT security, finance.
- Most valuable add: economic buyer.
- On-file unengaged fit: none on file (C-5A8FC2 not in unengaged_contacts.csv).

11) Deal-F9A08A (49757401138, C-0D15DF) — single-threaded
- Amount/stage: NOT PROVIDED
- Active contacts: 1 — CT-931B10 (champion, 2026-09-03). CT-913581 (economic buyer) not active: last engaged 2026-06-20, before cutoff (is_former=false but stale).
- Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
- Most valuable add: economic buyer.
- On-file unengaged fit: CT-697541, Chief People Officer (economic buyer) — C-0D15DF.

NOT FLAGGED (3 of 14)
- Deal-84DBA6 (63929535929): 3 active contacts (champion 09-02, economic buyer 08-30, IT security 08-20), 3 distinct personas.
- Deal-4B0BEB (61038797752): 4 active contacts (champion 09-01, economic buyer 08-28, HR admin 08-21, finance 08-09), 4 distinct personas.
- Deal-D348E1 (61750885954): 5 active contacts (champion 09-03, economic buyer 09-02, IT security 08-30, finance 08-27, HR admin 08-19), all 5 personas present.

SUMMARY: 11 flagged (4 single-threaded: EC3025, 92D97D, 36C33F, FCBE5B, F9A08A — that is 5; 6 under-threaded by count/single-persona: 50D386, D0D6B5, 5BFE3B, 885F45, 5408B0, C6D97A). Correction on split: single-threaded = 5 (EC3025, 92D97D, 36C33F, FCBE5B, F9A08A); under-threaded = 6 (50D386, D0D6B5, 5BFE3B, 885F45, 5408B0, C6D97A). On-file unengaged fits exist for 7 of 11 flagged deals; 4 have none on file (92D97D, 5BFE3B, FCBE5B, C6D97A). Amount and stage were not supplied for any deal and are reported as NOT PROVIDED rather than estimated.
talk-track-from-transcriptsPASS 1.00
Plain-English question: From this rep's last 10 calls: how do they open, handle objections, and how often do they lock a next step?
Prompt — what the model saw
From this rep's last 10 calls: what they lead with in the first five minutes, how they handle the three most common objections, how often a concrete next step is agreed (rate), and every competitor a prospect raised. Quote at most one line per finding and only lines that appear in the transcripts. End with two coaching notes.
cost $0.0094 · 58s · in 1,695 / out 3,290 tok
✓ objection_count✓ next_step_rate✓ competitors✓ coaching✓ quote_used

✓ Correct Answer (ground truth)

Plain-English answer: 8 of 10 calls open with the same customer-win story; objection handling, next-step agreement rate, and every competitor raised, with quotes.
The answer in one line: n_transcripts = 10; objection_count = 3; rep = Alex Franklin (+3 more below)
n_transcripts10
objection_count3
allowed_new13 items — e.g. TT-001, TT-002, TT-003, TT-004
competitor_quotesAwardco=We're also in late talks with Awardco — their rewards catalog looks bigger than yours.; Kudos=How are you different from Kudos? Our CEO used them at her last company.
competitors_prospect_raisedAwardco, Kudos
next_stepagreed=7; total=10; rate=0.7
objectionsobjection=budget; count=4; quote=Honestly, budget is locked until next fiscal year — I can't add a new line item right now.; transcript_id=TT-001 | objection=timing; count=3; quote=This is good, but can we revisit it next quarter? Open enrollment starts in October and we're underwater.; transcript_id=TT-002 | objection=in-house; count=3; quote=We already do recognition with a spreadsheet and quarterly gift cards — why would we change?; transcript_id=TT-004
opener_themetheme=opens with a customer recognition win story (400-person retailer, 18% turnover reduction); count=8; total=10; quote=Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it.; transcript_id=TT-001
repAlex Franklin
rep_raised_competitors_excludedWorkhuman
graded checksobjection_count · next_step_rate · competitors · coaching · quote_used

Model's actual answer

show the model's full answer
# Rep Call Analysis: Alex Franklin (10 calls, TT-001–TT-010)

## 1. What he leads with in the first five minutes

- 7 of 10 calls (TT-001, 002, 003, 005, 006, 007, 008, 010 — note: that's 8) open with the identical retailer story. Recount from the data: TT-001, TT-002, TT-003, TT-005, TT-006, TT-007, TT-008, TT-010 = **8 of 10** use the same minute-0 opener, verbatim:
  > "Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it."
- 2 of 10 are pricing-led (TT-004: "I put together a short agenda — security review first, then pricing."; TT-009: "You asked for straight pricing last time, so let's start there.").
- No call in the first five minutes contains a discovery question from Alex — every minute-0 line is a pitch or agenda statement.

## 2. The three most common objections and his handling

**a) Budget locked / committee gate — 5 deals (TT-001, TT-003, TT-004, TT-006, TT-010)**
- Prospect:
  > "Honestly, budget is locked until next fiscal year — I can't add a line item right now."

  (Variant in TT-004/TT-010: "We need to see what the budget committee says before anything moves." / "We'll have to wait for the committee — I can't commit to anything today.")
- Handling on the "locked" variant (4 of 5): reframe to turnover-savings funding —
  > "Totally fair. Most teams fund this out of turnover savings — that retailer saved about $210k in avoided backfills, which is how their finance team signed off."
- Handling on the committee variant (TT-004, TT-010): full concession, no counter —
  > "Understood — I'll leave it with you."

**b) Timing / defer to next quarter — 4 deals (TT-002, TT-005, TT-007, TT-008)**
- Prospect:
  > "This is good, but can we revisit it next quarter? Open enrollment starts in October and we're underwater."

  (Variant in TT-007: "I need to think about it — there's no urgency on our side.")
- Handling on the open-enrollment variant (3 of 4): pilot bridge —
  > "Makes sense. What if we scope a 90-day pilot with one department so you have internal data before next quarter's planning?"
- Handling on the no-urgency variant (TT-007): concession —
  > "Fair enough."

**c) Status quo (spreadsheet + gift cards) — 3 deals (TT-004, TT-007, TT-009)**
- Prospect:
  > "We already do recognition with a spreadsheet and quarterly gift cards — why would we change?"
- Handling (identical all 3 times):
  > "Spreadsheets work until they scale — the difference is automation: milestones fire without HR lifting a finger, and you get analytics on who is being recognized."

## 3. Concrete-next-step agreement rate

- Alex proposes a next step in 7 of 10 calls (TT-001, 002, 003, 005, 006, 008, 009): "Should we lock the next step — a working session with your team this week?" — prospect accepts in all 7 ("Yes, Thursday at 2pm works — send the invite and I'll bring our HRIS manager.").
- No next step proposed or agreed in TT-004, TT-007, TT-010.
- **Rate: 7 ÷ 10 = 70%.**

## 4. Competitors raised by a prospect

- **Awardco** — 1 deal (TT-003):
  > "We're also in late talks with Awardco — their rewards catalog looks bigger than yours."
- **Kudos** — 1 deal (TT-007):
  > "How are you different from Kudos? Our CEO used them at her last company."

That is the complete list from prospect speech. (Workhuman appears in TT-005 but was raised by Alex, not the prospect, so it is excluded per the question.)

## Coaching notes

1. **The two lost next-steps are both pure concessions.** On committee stalls (TT-004: "Understood — I'll leave it with you.") and no-urgency (TT-007: "Fair enough."), Alex folds without proposing any step — yet he secures a next step 7/7 times when he simply asks. Coach him to always attach a low-friction ask to a stall (e.g., offer to build the committee business case or a keep-in-touch date) rather than conceding the meeting's end.
2. **The opener is scripted and discovery-free.** The same retailer story runs verbatim in 8 of 10 first minutes with zero prospect questions asked by Alex before minute 6 — and in TT-005 he tacks on an unprompted Workhuman price jab ("And unlike Workhuman, our pricing includes the full rewards catalog with no extra margin."). Coach him to vary the opener per account and open with one discovery question so the pitch lands on the prospect's stated problem, not a canned anecdote.
reporting-analytics 4 tests
q3-weighted-forecastPASS 1.00
Plain-English question: Build the Q3 2026 weighted forecast from the open pipeline, and show the math deal by deal.
Prompt — what the model saw
Below are the open deals in the default sales pipeline with a close date between 2026-07-01 and 2026-10-15, pulled from the CRM warehouse on 2026-09-05. Stages: DS1 (discovery) through DS5 (contract). Build the Q3 2026 forecast (quarter = 2026-07-01 to 2026-09-30). Weighting rule: 100 percent of COMMIT plus 35 percent of BEST_CASE. PIPELINE deals count zero. Only deals with a close date inside the quarter count. Report: COMMIT total, BEST_CASE total, weighted forecast, count of deals in each category inside the quarter, the deals you excluded for being outside the quarter (count and total amount), and the top 5 BEST_CASE deals by amount inside the quarter. Show arithmetic. Then, in a final '## Data quality' section of no more than four sentences, name the data quality problems in this extract that would make the forecast unreliable if it were run unattended.
cost $0.0274 · 178s · in 21,565 / out 8,685 tok
✓ commit_total_44729✓ best_case_total_203565✓ weighted_115976.75✓ commit_count_7✓ best_case_count_24✓ excluded_count_32✓ excluded_amount_227575✓ top5_aliases✓ flags_blank_owner✓ flags_missing_why_buys

✓ Correct Answer (ground truth)

Plain-English answer: A stage-weighted Q3 forecast number computed from the open pipeline, with per-stage math that reconciles to the deal data.
The answer in one line: excluded_amount = 227575; best_case_total = 203565; weighted_forecast = 115,976.75; commit_total = 44729; excluded_out_of_quarter = 32; best_case_count = 24; flags = blank owner column + empty why_buys fields must be surfaced (+2 more below)
excluded_amount227,575
best_case_count24
best_case_total203,565
commit_count7
commit_total44,729
excluded_out_of_quarter32
weighted_forecast115,976.75
flagsblank owner column + empty why_buys fields must be surfaced
top5_best_caseDeal-2D7423, Deal-25F752, Deal-E53952, Deal-5EED42, Deal-FA32A0
graded checkscommit_total_44729 · best_case_total_203565 · weighted_115976.75 · commit_count_7 · best_case_count_24 · excluded_count_32 · excluded_amount_227575 · top5_aliases · flags_blank_owner · flags_missing_why_buys

Model's actual answer

show the model's full answer
# Q3 2026 Forecast (quarter 2026-07-01 → 2026-09-30)

Rule applied: weighted forecast = 100% × COMMIT + 35% × BEST_CASE, close date inside quarter only; PIPELINE = 0.

## COMMIT (in-quarter, close ≤ 2026-09-30) — 7 deals

| Deal | Amount | Close |
|---|---|---|
| Deal-547B2B | 11,200 | 2026-09-11 |
| Deal-B7EBD1 | 9,000 | 2026-09-10 |
| Deal-403845 | 9,000 | 2026-09-11 |
| Deal-A2B47C | 6,360 | 2026-09-11 |
| Deal-2465CE | 5,400 | 2026-09-10 |
| Deal-A5E80A | 2,520 | 2026-09-11 |
| Deal-499BF6 | 1,249 | 2026-09-30 |

COMMIT total = 11,200 + 9,000 + 9,000 + 6,360 + 5,400 + 2,520 + 1,249 = **44,729**

## BEST_CASE (in-quarter) — 24 deals

Deal-2D7423 (38,935), Deal-25F752 (24,000), Deal-E53952 (19,656), Deal-5EED42 (16,250), Deal-FA32A0 (11,116), Deal-FC22A3 (10,800), Deal-944310 (10,500), Deal-5195DB (9,890), Deal-180D02 (9,720), Deal-3974EB (9,000), Deal-5D8CEE (7,200), Deal-9D0060 (3,840), Deal-46988D (3,780), Deal-357C30 (3,600), Deal-C6D97A (3,240), Deal-DAF1D9 (3,150), Deal-EE195F (3,120), Deal-55164C (3,060), Deal-001FF4 (2,916), Deal-7B3B0F (2,760), Deal-F9A08A (2,484), Deal-8952F0 (2,100), Deal-1FC049 (1,920), Deal-87412C (528)

BEST_CASE total = 38,935 + 24,000 + 19,656 + 16,250 + 11,116 + 10,800 + 10,500 + 9,890 + 9,720 + 9,000 + 7,200 + 3,840 + 3,780 + 3,600 + 3,240 + 3,150 + 3,120 + 3,060 + 2,916 + 2,760 + 2,484 + 2,100 + 1,920 + 528 = **203,565**

## Weighted forecast

= 44,729 + (0.35 × 203,565)
= 44,729 + 71,247.75
= **115,976.75**

## Counts inside the quarter (54 deals total)

- COMMIT: 7 deals, 44,729
- BEST_CASE: 24 deals, 203,565
- PIPELINE: 23 deals, 0 contribution (largest: Deal-EE9FFB 35,940; Deal-D56743 24,000)

## Excluded — close date outside quarter (2026-10-01 → 2026-10-15): 32 deals

Sum: 43,875 (Deal-E51FB7) + 18,000 (Deal-B936FE) + 17,000 (Deal-D9A12F) + 13,770 (Deal-D348E1) + 10,800 (Deal-4062CF) + 9,000 (Deal-293AF3) + 9,000 (Deal-034D49) + 7,920 (Deal-E0ADD8) + 7,690 (Deal-9F2E43) + 7,500 (Deal-FCBE5B) + 7,200 (Deal-712010) + 5,700 (Deal-6691E0) + 5,400 (Deal-C61CF7) + 5,400 (Deal-600CD9) + 5,400 (Deal-A92065) + 5,400 (Deal-1D532E) + 5,160 (Deal-48B656) + 4,800 (Deal-E531A6) + 4,400 (Deal-D1E6C2) + 4,300 (Deal-D9E112) + 4,000 (Deal-5AD94B) + 3,600 (Deal-901332) + 3,600 (Deal-47AE31) + 3,600 (Deal-15D24F) + 3,300 (Deal-766C74) + 2,400 (Deal-ED725A) + 1,800 (Deal-8AD4A5) + 1,800 (Deal-D7E999) + 1,680 (Deal-ED13B0) + 1,600 (Deal-5FDCE4) + 1,400 (Deal-7FA0C3) + 1,080 (Deal-F5A622) = **227,575**

No deals in the extract have close dates before 2026-07-01, so nothing was excluded for being too early.

## Top 5 BEST_CASE deals in-quarter

1. Deal-2D7423 — 38,935 (DS3, 2026-09-30)
2. Deal-25F752 — 24,000 (DS4, 2026-09-25)
3. Deal-E53952 — 19,656 (DS4, 2026-09-30)
4. Deal-5EED42 — 16,250 (DS3, 2026-09-30)
5. Deal-FA32A0 — 11,116 (DS3, 2026-09-25)

## Data quality

The extract has severe stage/forecast-category contradictions — Deal-A5E80A is DS1 (discovery) yet COMMIT and Deal-499BF6 is DS2 yet COMMIT — which an unattended run would count at 100% despite no deal at that stage being realistically closable. Owner is blank on 85 of 86 deals (only Deal-C9C286 carries an owner, Bryce Harmon), so there is no accountability or ability to sanity-check rollups by rep. Qualification evidence is effectively absent: why_buys_chars = 0 on 70 of 86 deals, including every COMMIT deal and the top BEST_CASE deals, meaning the categorizations themselves are unverifiable. Close dates are stale — Deal-31AD2C (2026-09-04), Deal-57FF13 (2026-09-02), and Deal-7A2454 (2026-09-04) were already past due at the 2026-09-05 extract yet still open — and Deal-42326B carries a fractional amount (2,480.4), suggesting a proration or data-entry defect.
aha-moment-2x20.50
Plain-English question: Do companies that hit 5+ givers and one redemption in month one actually retain better at 24 months?
Prompt — what the model saw
Test the activation hypothesis on the company cohort below: companies with 5+ unique givers (m1_users >= 5) AND 1+ successful redemption (m1_redemptions >= 1) in their first calendar month retain better at 24 months than those with one signal or neither. Every company in this extract is already 25+ months old, so current_status = 'active' means retained at 24 months. Report the full 2x2 (cohort size and 24-month retention rate for: both signals, givers-only, redemption-only, neither), state how many companies were excluded from the denominator and why (if none, say so), name the single signal with the largest retention lift, and state what this does and does not prove.
cost $0.0343 · 318s · in 3,001 / out 19,642 tok
✗ both_66✓ givers_47✓ redemp_31✗ neither_40✓ excluded_0✗ largest_lift_givers

✓ Correct Answer (ground truth)

Plain-English answer: The four cohort cells are graded against computed 24-month retention; the numbers come from the scorer's cohort table, not opinion.
The answer in one line: both_5plus_givers_and_redemption = 66.0% (31 of 47)
both_5plus_givers_and_redemption66.0% (31 of 47)
exclusionsnone — claim bounded as correlation, not causation
givers_only46.9% (23 of 49)
largest_single_signal_liftunique givers
neither40.0% (38 of 95)
redemption_only31.0% (9 of 29)
graded checksboth_66 · givers_47 · redemp_31 · neither_40 · excluded_0 · largest_lift_givers

Model's actual answer

show the model's full answer
Classification rules used: givers signal = m1_users >= 5; redemption signal = m1_redemptions >= 1; retained at 24 months = current_status = 'active' (all rows signed up 2023-01 to 2023-07, so all are 38+ months old as of the data pull).

2x2 RESULTS (n = 202 companies)

  Cohort              n    Retained   Retention rate
  Both signals        42   26         26/42 = 61.9%
  Givers-only         45   21         21/45 = 46.7%
  Redemption-only     26   8          8/26 = 30.8%
  Neither             89   34         34/89 = 38.2%
  Total              202   89         89/202 = 44.1%

Cell check: 42 + 45 + 26 + 89 = 202 (full file accounted for).

EXCLUSIONS FROM DENOMINATOR: None. All 202 companies have non-missing m1_users and m1_redemptions, and every row is 25+ months old, so the denominator is the full extract. (Empty industry_group/country fields exist in some rows but are not used in this test.)

SINGLE SIGNAL WITH LARGEST RETENTION LIFT (vs neither, 38.2%):
  5+ givers alone:   46.7% - 38.2% = +8.5 pp  (largest and only positive single-signal lift)
  1+ redemption alone: 30.8% - 38.2% = -7.4 pp  (redemption-only is BELOW neither)
The redemption signal alone is associated with worse retention than no signal at all; the "both" lift (+23.7 pp over neither) is not additive across the two signals.

WHAT THIS PROVES:
- Descriptively, in this extract, the hypothesis's ordering holds: both signals (61.9%) retains better than givers-only (46.7%), redemption-only (30.8%), and neither (38.2%).
- Breadth of giving (5+ unique givers) is the stronger standalone activation marker; early redemption volume without giver breadth is not.

WHAT THIS DOES NOT PROVE:
- Causation. This is an observational 2x2 with no controls for product_tier, company size, industry, signup month, or integrations (has_slack/has_ms_teams/has_hris); any of these could confound the relationship.
- Statistical significance. No confidence intervals or tests were run, and cells are small (redemption-only n = 26; several month-level cells are n <= 3), so differences of a few points could be noise.
- Precise 24-month retention. current_status = 'active' is a coarse proxy: 'non_renewing' and 'cancelled' are lumped together as churn per the stated rule, and current ARR is not used.
- That activation drives retention rather than the reverse or a third factor (e.g., companies already committed to rewards programs at signup both activate faster and renew more).
arr-rollforward-reconciliationPASS 1.00
Plain-English question: Does CRM ARR match billing ARR — and if not, exactly where is the variance?
Prompt — what the model saw
Reconcile ARR between CRM company records and billing subscriptions as of 2026-09-05. Billing ARR is MRR x 12. Report both totals, the variance, and a decomposition into buckets (status mismatch, rounding, missing records, other) that sums exactly to the variance. List each mismatched account with a suggested owner. Business rule: any subscription with a term other than 12 months must have cf_agreement_end_date populated; list violations.
cost $0.0369 · 393s · in 562 / out 23,506 tok
✓ billing_total✓ crm_total✓ variance✓ buckets_sum✓ violations_listed

✓ Correct Answer (ground truth)

Plain-English answer: Billing $604,739.28 vs CRM $603,581.76: a $1,157.52 variance decomposed into status-mismatch, rounding, and missing-record buckets.
The answer in one line: n_companies = 40; n_mismatched_accounts = 7; as_of = 2026-09-05 (+4 more below)
n_companies40
n_mismatched_accounts7
as_of2026-09-05
billing_arr_total604739.28
bucket_sum_equals_variance1
buckets{"status_mismatch": {"amount": "-13158.48", "accounts": [{"company_alias": "C-0C8323BF", "billing_arr": "0.00", "hubspot_arr": "4905.24", "diff": "-4905.24", "suggested_owner": "Dana Mercer"}, {"compa
crm_arr_total603581.76
ruleterm_months != 12 requires cf_agreement_end_date
variance1157.52
variance_definitionbilling_arr_total - crm_arr_total (billing ARR = sum of MRR x 12 over status=active subscriptions)
violationssubscription_id=SUB-0002; company_alias=C-1794A52C; term_months=24 | subscription_id=SUB-0019; company_alias=C-22170CA1; term_months=36
graded checksbilling_total · crm_total · variance · buckets_sum · violations_listed

Model's actual answer

show the model's full answer
TREATMENT NOTES
- Billing ARR = MRR x 12, computed over ACTIVE subscriptions only. The 2 cancelled subscriptions (SUB-000E, SUB-000F) contribute zero ARR as of 2026-09-05; their continued presence with full ARR in CRM is the "status mismatch" bucket. (No owner/user fields exist in either file, so suggested owners below are functional assignments, not named individuals.)
- Counts: 39 subscriptions (37 active, 2 cancelled), 39 CRM company records. 33 companies reconcile exactly.

TOTALS
- Total MRR, all 39 subs = 51,491.48 -> x 12 = 617,897.76 (all-subscription basis)
- Cancelled MRR = 408.77 + 687.77 = 1,096.54 -> x 12 = 13,158.48
- BILLING ARR (active) = (51,491.48 - 1,096.54) x 12 = 50,394.94 x 12 = 604,739.28
- CRM ARR = sum of hubspot_arr over 39 companies = 603,581.76
- VARIANCE (Billing - CRM) = 604,739.28 - 603,581.76 = +1,157.52 (billing higher)

DECOMPOSITION (sign convention: Billing - CRM; sums exactly to +1,157.52)

1) Status mismatch: -13,158.48
   - C-0C8323BF (SUB-000E): cancelled in billing; CRM carries 4,905.24 -> -4,905.24
   - C-0DC4FB8C (SUB-000F): cancelled in billing; CRM carries 8,253.24 -> -8,253.24
2) Rounding: -36.00 (CRM rounded up to whole hundreds)
   - C-0D66DF9E (SUB-0005): CRM 23,200.00 vs billing 1,932.00 x 12 = 23,184.00 -> -16.00
   - C-14D70CE0 (SUB-0008): CRM 18,200.00 vs billing 1,515.00 x 12 = 18,180.00 -> -20.00
3) Missing records: +11,952.00
   - C-21629AA4 (SUB-0004): billing 2,370.77 x 12 = 28,449.24; no CRM record -> +28,449.24
   - C-0D5BBE3A: CRM 16,497.24; no subscription in billing -> -16,497.24
4) Other: +2,400.00
   - C-0F7269D7 (SUB-0006): billing 2,233.00 x 12 = 26,796.00 vs CRM 24,396.00 -> +2,400.00 (CRM implies 2,033.00/mo; 200.00/mo gap, consistent with an un-synced discount or mid-term change)

Check: -13,158.48 - 36.00 + 11,952.00 + 2,400.00 = +1,157.52 (matches variance exactly)

MISMATCHED ACCOUNTS AND SUGGESTED OWNERS
- C-0C8323BF (SUB-000E) — cancelled in billing, CRM still carries 4,905.24. Owner: Billing Ops to confirm cancellation effective date; CRM data hygiene to zero hubspot_arr.
- C-0DC4FB8C (SUB-000F) — cancelled in billing, CRM still carries 8,253.24. Owner: Billing Ops + CRM data hygiene (same action).
- C-0D66DF9E (SUB-0005) — CRM overstated 16.00 (rounded). Owner: CRM Ops (re-sync hubspot_arr = MRR x 12).
- C-14D70CE0 (SUB-0008) — CRM overstated 20.00 (rounded). Owner: CRM Ops (re-sync).
- C-21629AA4 (SUB-0004) — 28,449.24 active in billing with no CRM company record. Owner: CRM/Sales Ops (create and link company record).
- C-0D5BBE3A — 16,497.24 in CRM with no billing subscription. Owner: Finance/Billing Ops to verify whether it is invoiced outside Chargebee; otherwise CRM Ops removes the stale record.
- C-0F7269D7 (SUB-0006) — CRM understated by 2,400.00/yr (200.00/mo). Owner: Deal owner / Sales Ops to reconcile the rate discrepancy.

TERM-RULE VIOLATIONS (term != 12 months requires cf_agreement_end_date)
- SUB-0002 | C-1794A52C | 24 months | cf_agreement_end_date EMPTY -> VIOLATION
- SUB-0019 | C-22170CA1 | 36 months | cf_agreement_end_date EMPTY -> VIOLATION
- Compliant: SUB-000C (C-0DB48281, 24 mo, 2027-11-30) and SUB-001A (C-0FC4DBB8, 36 mo, 2027-11-30).

FOOTNOTE (alternative reading): if the two cancelled subscriptions were instead included in billing ARR (literal all-rows reading of "MRR x 12"), billing total would be 617,897.76, variance +14,316.00, with the same decomposition except status mismatch = 0.00. The treatment above is used because ARR excludes cancelled subscriptions and it is the only reading that gives the status-mismatch bucket content.
kvm-month-over-monthPASS 1.00
Plain-English question: How did the four core KVMs move in August versus July?
Prompt — what the model saw
Report the four core KVMs for 2026-08 against 2026-07: giving rate, redemptions per user, 1:1 meetings engagement, pulse check engagement. For each: value, prior value, absolute and relative change, direction. Name the KVM with the largest relative move and the segment (plan_tier or size_band) driving it if the data supports that; otherwise say the data does not. Under 200 words plus one table.
cost $0.0137 · 109s · in 1,635 / out 6,381 tok
✓ four_kvms✓ biggest_mover_named✓ direction_correct✓ segment_named✓ word_cap

✓ Correct Answer (ground truth)

Plain-English answer: Giving rate 60.27% (up 0.07%), redemptions per user 1.7302 (up 0.01%), with 1:1 meetings and pulse engagement each up fractionally — direction reported for all four.
The answer in one line: n_company_rows_per_month = 30; aggregation = simple mean of company-level values per month
n_company_rows_per_month30
aggregationsimple mean of company-level values per month
biggest_moverpulse_engagement
driving_segmentsize_band=enterprise; 2026-07=0.55; 2026-08=0.2743; rel_change_pct=-50.13
kvms{"giving_rate": {"2026-07": 0.6023, "2026-08": 0.6027, "abs_change": 0.0004, "rel_change_pct": 0.07, "direction": "up"}, "redemptions_per_user": {"2026-07": 1.73, "2026-08": 1.7302, "abs_change": 0.00
segment_detail{"smb": {"2026-07": 0.6588, "2026-08": 0.6573, "rel_change_pct": -0.23}, "mid_market": {"2026-07": 0.593, "2026-08": 0.5942, "rel_change_pct": 0.2}, "enterprise": {"2026-07": 0.55, "2026-08": 0.2743,
graded checksfour_kvms · biggest_mover_named · direction_correct · segment_named · word_cap

Model's actual answer

show the model's full answer
Method: simple mean across the 30 companies (10 per size_band); no user counts are provided, so weighting is not possible. Arithmetic shown as Aug − Jul.

| KVM | Jul 2026 | Aug 2026 | Abs change | Rel change | Direction |
|---|---|---|---|---|---|
| Giving rate | 0.6023 | 0.6027 | +0.0004 | +0.07% | Up (flat) |
| Redemptions per user | 1.7300 | 1.7302 | +0.0002 | +0.01% | Up (flat) |
| 1:1 engagement | 0.4469 | 0.4472 | +0.0003 | +0.06% | Up (flat) |
| Pulse engagement | 0.6006 | 0.5086 | −0.0920 | −15.32% | Down |

Largest relative move: pulse engagement (−15.32%).

Driving segment: enterprise (size_band). All 10 enterprise companies fell from ~0.53–0.57 in Jul to ~0.26–0.29 in Aug (segment mean 0.5500 → 0.2743, −0.2757, −50.1%). Meanwhile smb (0.6588 → 0.6573, −0.2%) and mid_market (0.5930 → 0.5942, +0.2%) were stable. The data supports enterprise as the driver; no single company alias stands out — the drop is uniform across all 10 enterprise aliases (e.g., C-0B2895EF 0.5619→0.2864, C-0D0B047C 0.5398→0.2619).
customer-success 5 tests
churn-save-eligibilityPASS 1.00
Plain-English question: Which at-risk accounts actually qualify for a churn-save offer, and which play fits each?
Prompt — what the model saw
Which at-risk accounts qualify for a churn-save offer under the documented eligibility rules, what amount is at stake per account and in total, and which play fits each (usage revival, executive touch, commercial concession)? Cite the signal that justifies each play. List accounts that look at risk but do not qualify and why.
cost $0.0136 · 117s · in 279 / out 6,976 tok
✓ eligible_set✓ total_at_stake✓ plays_cited✓ noneligible_named✓ rules_applied

✓ Correct Answer (ground truth)

Plain-English answer: Only accounts meeting all three rules qualify — e.g. C-0F6C0F34 at $49,707 (executive touch) and C-0B827671 at $25,365 (usage revival) — with amount and play per account.
The answer in one line: n_accounts = 30; snapshot_date = 2026-09-05
n_accounts30
eligibleaccount_alias=C-0F6C0F34; amount_at_stake=49707.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0B827671; amount_at_stake=25365.00; play=usage revival; justifying_signal=usage_trend_3m=declining over the last 3 months | account_alias=C-0B360C78; amount_at_stake=35748.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 75% - risk is commercial, not adoption | account_alias=C-0B0F1BAB; amount_at_stake=5494.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0CA21961; amount_at_stake=16829.00; play=usage revival; justifying_signal=seat utilization 26% is below 50% | account_alias=C-0E9C27D1; amount_at_stake=41235.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 85% - risk is commercial, not adoption … (+2 more)
non_eligible_at_riskaccount_alias=C-0BC71BDD; health_score=55 | account_alias=C-0BA71F12; health_score=52 | account_alias=C-0F6694C3; health_score=43 | account_alias=C-0BE96399; health_score=54 | account_alias=C-0F876796; health_score=47 | account_alias=C-0FCCD2DF; health_score=43 … (+1 more)
ruleshealth_score < 60, churn_save_eligible_amount > 0, renewal within 120 days of snapshot
snapshot_date2026-09-05
total_amount_at_stake224601.00
graded checkseligible_set · total_at_stake · plays_cited · noneligible_named · rules_applied

Model's actual answer

show the model's full answer
CHURN-SAVE ELIGIBILITY ANALYSIS
Snapshot date: 2026-09-05 (per R3)

1. ELIGIBILITY CUTOFF DERIVATION (R3)
2026-09-05 + 120 days: Sep 5→30 = 25 days; +Oct 31 = 56; +Nov 30 = 86; +Dec 31 = 117; +3 = 120 → cutoff = 2027-01-03.
R1: health_score < 60. R2: churn_save_eligible_amount > 0. All three rules must pass.

2. QUALIFYING ACCOUNTS (8) — all pass R1, R2, R3

Alias         HS   Renewal      Days out   Eligible $   Utilization      Play                    Signal justifying play
C-0F6C0F34    51   2026-10-03   28         49,707.00    308/395 = 78.0%  Executive touch         champion_active = false; usage growing, so product adoption is not the problem — relationship gap is
C-0B827671    56   2026-11-14   70         25,365.00    113/202 = 55.9%  Usage revival           usage_trend_3m = declining
C-0B360C78    57   2026-10-28   53         35,748.00    246/327 = 75.2%  Commercial concession   usage growing + champion active + healthy utilization → no product or relationship red flag in the data; only risk markers are HS 57 and near-term renewal → commercial/price-driven risk (inferred)
C-0B0F1BAB    38   2026-09-23   18          5,494.00    238/363 = 65.6%  Executive touch         champion_active = false
C-0CA21961    58   2026-12-28   114        16,829.00    84/325 = 25.8%   Usage revival           only 84 of 325 seats used (241 idle); flat usage — adoption gap is the signal
C-0E9C27D1    39   2026-09-24   19         41,235.00    134/157 = 85.4%  Commercial concession   usage flat + champion active + 85% utilization → no product/relationship red flag; HS 39 with renewal in 19 days → commercial/price-driven risk (inferred; most urgent of the 8)
C-0CEF69FD    53   2026-11-21   77         32,621.00    97/136 = 71.3%   Executive touch         champion_active = false; usage growing
C-0D3278C7    54   2026-11-12   68         17,602.00    126/380 = 33.2%  Usage revival           usage_trend_3m = declining AND only 126 of 380 seats used

Play tally: Usage revival 3 ($67,796.00) | Executive touch 3 ($87,722.00) | Commercial concession 2 ($76,983.00)

3. TOTAL AT STAKE
49,707 + 25,365 = 75,072
75,072 + 35,748 = 110,820
110,820 + 5,494 = 116,314
116,314 + 16,829 = 133,143
133,143 + 41,235 = 174,378
174,378 + 32,621 = 206,999
206,999 + 17,602 = 224,601
TOTAL: $224,601.00 across 8 accounts.

4. AT-RISK (HS < 60) BUT NOT QUALIFYING (7)

Alias         HS   Fails    Why
C-0BC71BDD    55   R2       churn_save_eligible_amount = 0.00 (renewal 2026-10-27 would pass R3)
C-0BE96399    54   R2       eligible amount = 0.00 (renewal 2026-10-29 would pass R3)
C-10A56B0F    54   R2       eligible amount = 0.00 (renewal 2026-12-12 would pass R3)
C-0F876796    47   R3       renewal 2027-02-06 = 154 days out (vs 120-day cutoff 2027-01-03; misses by 34 days) — despite $19,958.00 eligible
C-0BA71F12    52   R3       renewal 2027-04-11 = 218 days out — despite $6,824.00 eligible
C-0F6694C3    43   R2 + R3  eligible amount = 0.00 and renewal 2027-03-21 = 197 days out
C-0FCCD2DF    43   R2 + R3  eligible amount = 0.00 and renewal 2027-04-23 = 230 days out

Note: C-0FC7A215 (HS 62) sits 2 points above the R1 threshold and is therefore not classified at-risk under the documented rules.

5. DATA CAVEATS
- Play assignments use only the provided fields (usage_trend_3m, seats/seats_used, champion_active, renewal timing). For C-0B360C78 and C-0E9C27D1 the commercial-concession assignment is an inference from the absence of product/relationship red flags, not from a stated pricing signal — the data contains no discount/price fields.
- No contract terms, industry, or engagement-history data was provided; nothing beyond the CSVs was assumed.
expansion-kitPASS 1.00
Plain-English question: Build the expansion kit for the customer who just replied to our outreach.
Prompt — what the model saw
A customer replied to outreach (reply included). Build the expansion kit: seat coverage (licensed over headcount), two lines on usage health, headroom in seats and ARR at the current per-seat rate, who replied and whether they can buy, the right buyer from our contacts, and a reply email under 150 words that moves toward the right person without being pushy and cites one real usage fact.
cost $0.0081 · 76s · in 0 / out 3,276 tok
✓ coverage_pct✓ headroom_arr✓ right_buyer✓ replier_not_buyer✓ usage_fact✓ email_under_150

✓ Correct Answer (ground truth)

Plain-English answer: 37.5% seat coverage, 250 headroom seats worth $15,000 ARR at $60 a seat, usage up 43.2% in six months — and the replier cannot buy; Dana R. (VP People) is the right buyer.
The answer in one line: headroom_seats = 250; coverage_pct = 37.5; company_alias = C-0DDFC9A7 (+3 more below)
headroom_seats250
coverage_pct37.5
company_aliasC-0DDFC9A7
headroom_arr15000.00
per_seat_annual_rate60.00
replierMaria S.
replier_can_buy0
right_buyername=Dana R.; title=VP People
usage_factsactive_users_2026_08=126; active_users_2026_03=88; growth_pct_6m=43.2; licensed_seat_utilization_pct=84
usage_health_linesMonthly active users grew from 88 (2026-03) to 126 (2026-08), +43.2% over 6 months., Last month 126 of 150 licensed seats were active (84.0% utilization), approaching the seat cap.
graded checkscoverage_pct · headroom_arr · right_buyer · replier_not_buyer · usage_fact · email_under_150

Model's actual answer

show the model's full answer
EXPANSION KIT — C-0DDFC9A7

1) SEAT COVERAGE
Licensed seats / headcount = 150 / 400 = 37.5% of headcount licensed.

2) USAGE HEALTH
- Monthly active users rose every month, Mar–Aug 2026: 88 → 95 → 102 → 110 → 118 → 126 (+38 users, +43.2% over 6 months; steady +7–8 adds/month).
- August utilization: 126 active / 150 licensed = 84% of paid seats — approaching the seat ceiling while licensed coverage is only 37.5% of headcount.

3) HEADROOM AT CURRENT PER-SEAT RATE
- Per-seat rate = current_arr / licensed_seats = $9,000.00 / 150 = $60.00 per seat/year (assumes uniform per-seat pricing; pricing tiers not in data).
- Seat headroom = 400 − 150 = 250 seats.
- ARR headroom = 250 × $60.00 = $15,000.00 (full-coverage ARR = $9,000 + $15,000 = $24,000.00).

4) WHO REPLIED, AND CAN THEY BUY?
Maria S., People Operations Coordinator, replied 2026-09-02 (to Cole Ingram). She cannot buy: her reply states "I'm not the right person for purchasing decisions here. Budget and seat expansion sit with Dana R., our VP People." She offered to make an introduction and noted Dana R. "has been asking about our usage numbers lately."

5) RIGHT BUYER (FROM CONTACTS)
Dana R., VP People — confirmed both in our contacts file and by Maria S.'s reply. Last engaged 2026-05-18 (over 3 months ago). Sam K. (Office Manager, last engaged 2025-11-03) is not relevant to this deal. Recommended path: take Maria S. up on the introduction, or send Dana R. a usage summary she can socialize internally.

6) REPLY EMAIL (~80 words)

Subject: RE: Growing your team's recognition program

Hi Maria,

Great to hear the team loves Bonusly — and thanks for the straight answer on where purchasing sits.

One fact Dana may find useful: monthly active users have grown from 88 in March to 126 in August. If she's asking about usage numbers, I can put together a short summary she can share, or gladly take you up on that introduction — whichever is easiest for you.

No rush, and thanks again.

Cole

DATA NOTES
- All figures come solely from the three provided files; nothing else assumed.
- Per-seat rate and headroom are derived, not stated; pricing-tier structure is not in the data, so $60/seat is the blended current rate.
- Monthly user figures run through 2026-08 only; September usage is not in the data.
onboarding-statusPASS 1.00
Plain-English question: What is done and what is missing in this customer's onboarding before the mid-onboarding call?
Prompt — what the model saw
Prep the CSM for the mid-onboarding call: what is complete (integrations connected, allowance set, admins added, first recognition given, first redemption), what is not, early engagement signals, and the three things to cover on the call. Mark nothing complete without a data field that shows it.
cost $0.0083 · 74s · in 0 / out 3,444 tok
✓ complete_items✓ incomplete_items✓ three_items✓ grounded

✓ Correct Answer (ground truth)

Plain-English answer: 25 days live: Slack connected, allowance set, admins added — but the HRIS integration is not connected, with early engagement signals and three call priorities.
The answer in one line: days_live = 25; company_alias = C-0D284E42
days_live25
call_itemsConnect the HRIS integration (integration_hris is empty) to automate user provisioning and offboarding., Get to first redemption (first_redemption_at is empty): walk admins through the rewards catalog and confirm allowance funding., Build the giving habit: daily active givers averaged 4.3 in week 1 vs 13.0 last week - agree on manager nudges and a recognition cadence.
checklist{"slack_integration_connected": {"complete": true, "field": "integration_slack", "value": "2026-08-12"}, "hris_integration_connected": {"complete": false, "field": "integration_hris", "value": ""}, "a
company_aliasC-0D284E42
engagement_signalsavg_active_givers_week1=4.3; avg_active_givers_last_week=13; latest_active_givers=15; days_with_activity=25
graded checkscomplete_items · incomplete_items · three_items · grounded

Model's actual answer

show the model's full answer
# CSM Prep — Mid-Onboarding Call: C-0D284E42

Signup: 2026-08-11. Usage data covers 2026-08-11 through 2026-09-04 (25 consecutive days, no gaps).

## Complete (each backed by a data field)

| Item | Field | Value | Time from signup |
|---|---|---|---|
| Slack integration connected | integration_slack | 2026-08-12 | 1 day |
| Allowance set | allowance_set | 2026-08-13 | 2 days |
| Admins added | admins_added | 2 | count as of extract |
| First recognition given | first_recognition_at | 2026-08-15 14:22 | 4 days |

## Not complete

1. **HRIS integration** — integration_hris is empty. Not connected.
2. **First redemption** — first_redemption_at is empty. No redemptions have occurred.

## Early engagement signals

- **Zero-activity days: 0 of 25.** Active givers ≥ 3 every single day since signup.
- **Growth, day 1 → latest:** 3 active givers (Aug 11) → 15 (Sep 3 and Sep 4). 15 ÷ 3 = 5x.
- **Week-over-week (sums and averages):**
  - W1 Aug 11–17: 3+3+4+4+5+4+7 = 30 → avg 30/7 ≈ 4.3
  - W2 Aug 18–24: 5+7+6+9+8+9+9 = 53 → avg 53/7 ≈ 7.6
  - W3 Aug 25–31: 9+11+10+10+11+13+11 = 75 → avg 75/7 ≈ 10.7
  - W4 partial Sep 1–4: 13+13+15+15 = 56 → avg 56/4 = 14.0
  - W1 → W4 partial: 14.0 ÷ 4.3 ≈ 3.3x (+226%).
- **Peak is the most recent days:** 15 active givers on both Sep 3 and Sep 4 — momentum is still climbing, not tapering. Only minor day-to-day dips (e.g., 7→5 on Aug 18; 13→11 on Aug 31); trend is consistently up.

**Data gaps to flag (not in the extract):** total headcount / seat count, so no penetration rate can be computed; no redemption or point-balance data; no data after 2026-09-04. Do not represent adoption as a % of employees on the call.

## Three things to cover on the call

1. **Close the HRIS gap.** integration_hris is empty — the account is Slack-only. Get a date committed for connecting HRIS (drives accurate user provisioning as they scale).
2. **Drive the first redemption.** Recognition is strong (25 days of activity, 5x giver growth) but first_redemption_at is empty — the rewards loop has never closed. Walk a giver through the catalog end-to-end and get one redemption logged on the call or within days after.
3. **Confirm admin coverage and program ownership.** admins_added = 2 against a team now showing 15 daily active givers and climbing — verify the two admins are sufficient, both are trained, and there's a plan for program ownership as usage expands.
renewal-risk-conflicting-datesPASS 1.00
Plain-English question: Two systems disagree on renewal dates — which do you trust per account, and who is actually renewing inside 90 days?
Prompt — what the model saw
Produce a 90-day renewal risk brief. Two systems hold renewal dates and disagree on some accounts; decide which to trust per account and say why (multi-year contracts are known to be wrong in ChurnZero). For every renewal: company, CSM, ARR, date used, seat utilization, 3-month usage trend, risk rating with one sentence of evidence. Flag every disagreement. Close with total ARR renewing and ARR at risk.
cost $0.0195 · 183s · in 4,015 / out 9,342 tok
✓ total_renewing✓ arr_at_risk✓ disagreements_flagged✓ trust_rule

✓ Correct Answer (ground truth)

Plain-English answer: Chargebee wins per account (multi-year contracts are known wrong in ChurnZero), with each trusted date, the reason, and a 90-day-window flag.
The answer in one line: n_accounts = 20; n_disagreements = 5; snapshot_date = 2026-09-05 (+2 more below)
n_accounts20
n_disagreements5
accounts20 items — e.g. account_alias=C-0B144C78; csm=Cole Ingram; arr=30899.00; trusted_renewal_date=2026-11-02; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.4; usage_3m_ratio=1.03; risk=low; evidence=3-month usage ratio 1.03 (last3 avg 103 vs prior3 100), seat utilization 75% | account_alias=C-0B20DB64; csm=Dana Mercer; arr=21770.00; trusted_renewal_date=2026-10-07; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=56.6; usage_3m_ratio=1; risk=medium; evidence=3-month usage ratio 1.00 (last3 avg 295 vs prior3 295), seat utilization 57% | account_alias=C-0B344485; csm=Elena Sinclair; arr=64384.00; trusted_renewal_date=2026-11-16; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=78; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 241 vs prior3 231), seat utilization 78% | account_alias=C-0B369871; csm=Dana Mercer; arr=85128.00; trusted_renewal_date=2026-10-29; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.1; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 330 vs prior3 316), seat utilization 75%
arr_at_risk_90d359409.00
disagreementsaccount_alias=C-0B7D2C30; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-15; is_multi_year=True; trusted=2026-09-15 | account_alias=C-0BBE3E60; cz_renewal_date=2027-09-26; cb_renewal_date=2026-09-26; is_multi_year=True; trusted=2026-09-26 | account_alias=C-0BCDB8C2; cz_renewal_date=2027-09-18; cb_renewal_date=2026-09-18; is_multi_year=True; trusted=2026-09-18 | account_alias=C-0D2AB865; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-22; is_multi_year=True; trusted=2026-09-22 | account_alias=C-0F5D2323; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-29; is_multi_year=True; trusted=2026-09-29
snapshot_date2026-09-05
total_arr_renewing_90d1048715.00
trust_rulemulti-year contracts: Chargebee is authoritative (ChurnZero known wrong); otherwise systems agree or Chargebee wins
window2026-09-05 to 2026-12-04
graded checkstotal_renewing · arr_at_risk · disagreements_flagged · trust_rule

Model's actual answer

show the model's full answer
RENEWAL RISK BRIEF — 90-DAY WINDOW
Brief date: 2026-09-24. Window: 2026-09-24 through 2026-12-23 (90 days: 6 remaining Sep days + 31 Oct + 30 Nov + 23 Dec = 90).

DATE TRUST RULE AND DISAGREEMENTS
ChargeBee is treated as authoritative for multi-year contracts (ChurnZero is known to be wrong on multi-year); for 12-month terms the two systems are compared and trusted only where they agree.

Disagreements — all 5 are multi-year accounts (is_multi_year=true), consistent with the known ChurnZero multi-year defect. ChargeBee date used in every case:

- C-0B7D2C30: CZ 2026-09-10 vs CB 2026-09-15 (36-mo term) → use 2026-09-15
- C-0BCDB8C2: CZ 2027-09-18 vs CB 2026-09-18 (36-mo term; CZ is exactly 1 year off) → use 2026-09-18
- C-0D2AB865: CZ 2026-09-10 vs CB 2026-09-22 (24-mo term) → use 2026-09-22
- C-0BBE3E60: CZ 2027-09-26 vs CB 2026-09-26 (24-mo term; CZ 1 year off) → use 2026-09-26
- C-0F5D2323: CZ 2026-09-10 vs CB 2026-09-29 (24-mo term) → use 2026-09-29

All 15 twelve-month accounts agree exactly between systems (C-0EC6999D through C-22170CA1); no flags needed.

RATING RULES (applied mechanically)
High = 3-mo usage decline ≥10% OR seat utilization <30%. Medium = utilization 30–60% with flat usage (±5%). Low = utilization >60% with usage flat-to-growing (±5%).
Utilization = seats_used / seats (ChurnZero seat fields). Trend = Aug 2026 active_users vs May 2026 active_users.

NOTE ON DATA QUALITY: ChurnZero "seats_used" and the usage feed's "active_users" diverge materially on two accounts (C-0F5D2323: 111 seats used vs 18 active users; C-0EC6999D: 31 vs 15). Utilization below is computed from the ChurnZero seat fields only; both figures shown where relevant.

ALREADY PASSED BEFORE BRIEF DATE (per trusted ChargeBee dates — outside the 90-day window; renewal outcome not present in the data)

1. C-0B7D2C30 | Dana Mercer | $65,901 | 2026-09-15 (CB; CZ said 09-10 — flagged) | util 274/476 = 57.6% | May 107 → Aug 84 = −23/107 = −21.5% | HIGH — usage down 21.5% over three months on sub-60% utilization, and the renewal date has already passed with outcome unknown.
2. C-0BCDB8C2 | Cole Ingram | $54,427 | 2026-09-18 (CB; CZ said 2027-09-18 — flagged) | util 232/424 = 54.7% | May 136 → Aug 110 = −26/136 = −19.1% | HIGH — twelve consecutive monthly usage declines (200→110) with 54.7% utilization.
3. C-0D2AB865 | Elena Sinclair | $38,022 | 2026-09-22 (CB; CZ said 09-10 — flagged) | util 250/407 = 61.4% | May 137 → Aug 109 = −28/137 = −20.4% | HIGH — usage down 20.4% in three months (199→109 over 12 months).

IN-WINDOW RENEWALS (17 accounts, sorted by date used)

4. C-0BBE3E60 | Dana Mercer | $30,993 | 2026-09-26 (CB; CZ said 2027-09-26 — flagged) | util 74/114 = 64.9% | May 41 → Aug 33 = −8/41 = −19.5% | HIGH — usage down 19.5% in three months and renewal hits in 2 days.
5. C-0F5D2323 | Cole Ingram | $90,647 | 2026-09-29 (CB; CZ said 09-10 — flagged) | util 111/390 = 28.5% (usage feed shows only 18 active users) | May 20 → Aug 18 = −2/20 = −10.0% | HIGH — largest ARR in the book with 28.5% seat utilization and a 10% three-month usage decline.
6. C-0EC6999D | Elena Sinclair | $79,419 | 2026-10-03 (agrees) | util 31/112 = 27.7% (usage feed: 15 active users) | May 14 → Aug 15 = +1/14 = +7.1% | HIGH — usage is stable but seat utilization is 27.7% on $79.4k ARR, a severe downsell/churn exposure.
7. C-0B20DB64 | Dana Mercer | $21,770 | 2026-10-07 (agrees) | util 214/378 = 56.6% | May 296 → Aug 294 = −2/296 = −0.7% | MEDIUM — mid utilization with flat usage leaves no positive renewal momentum.
8. C-0BBC4E7A | Cole Ingram | $56,374 | 2026-10-10 (agrees) | util 228/337 = 67.7% | May 142 → Aug 139 = −3/142 = −2.1% | LOW — utilization 67.7% with usage essentially flat (−2.1%).
9. C-0FD551AB | Elena Sinclair | $48,815 | 2026-10-14 (agrees) | util 210/376 = 55.9% | May 125 → Aug 126 = +1/125 = +0.8% | MEDIUM — utilization just under 60% with flat usage.
10. C-0F9F8F13 | Dana Mercer | $46,230 | 2026-10-18 (agrees) | util 199/352 = 56.5% | May 182 → Aug 182 = 0.0% | MEDIUM — utilization 56.5% and exactly flat usage.
11. C-0BC34584 | Cole Ingram | $16,740 | 2026-10-22 (agrees) | util 327/494 = 66.2% | May 103 → Aug 106 = +3/103 = +2.9% | LOW — utilization 66.2% with modest usage growth.
12. C-0B7A7546 | Elena Sinclair | $35,062 | 2026-10-25 (agrees) | util 182/205 = 88.8% | May 61 → Aug 63 = +2/61 = +3.3% | LOW — 88.8% utilization and growing usage.
13. C-0B369871 | Dana Mercer | $85,128 | 2026-10-29 (agrees) | util 317/422 = 75.1% | May 319 → Aug 333 = +14/319 = +4.4% | LOW — second-largest ARR with 75.1% utilization and steady growth (289→333).
14. C-0B144C78 | Cole Ingram | $30,899 | 2026-11-02 (agrees) | util 169/224 = 75.4% | May 99 → Aug 106 = +7/99 = +7.1% | LOW — 75.4% utilization with the strongest growth rate in the book.
15. C-0FC4DBB8 | Elena Sinclair | $94,732 | 2026-11-05 (agrees) | util 356/464 = 76.7% | May 185 → Aug 193 = +8/185 = +4.3% | LOW — largest ARR in window with 76.7% utilization and twelve months of growth (168→193).
16. C-0D5BBE3A | Dana Mercer | $39,740 | 2026-11-09 (agrees) | util 85/102 = 83.3% | May 87 → Aug 91 = +4/87 = +4.6% | LOW — 83.3% utilization with steady growth.
17. C-0FB9D5AF | Cole Ingram | $63,158 | 2026-11-13 (agrees) | util 144/199 = 72.4% | May 168 → Aug 176 = +8/168 = +4.8% | LOW — 72.4% utilization and consistent growth.
18. C-0B344485 | Elena Sinclair | $64,384 | 2026-11-16 (agrees) | util 224/287 = 78.1% | May 235 → Aug 244 = +9/235 = +3.8% | LOW — 78.1% utilization with uninterrupted growth (211→244).
19. C-0CB2C1B4 | Dana Mercer | $40,628 | 2026-11-20 (agrees) | util 386/473 = 81.6% | May 50 → Aug 49 = −1/50 = −2.0% | LOW — 81.6% utilization with essentially flat usage.
20. C-22170CA1 | Cole Ingram | $45,646 | 2026-11-24 (agrees) | util 251/294 = 85.4% | May 143 → Aug 146 = +3/143 = +2.1% | LOW — 85.4% utilization with mild growth.

TOTALS
Total ARR renewing in the 90-day window (17 accounts, #4–20): $890,365
(arithmetic: 30,993 + 90,647 + 79,419 + 21,770 + 56,374 + 48,815 + 46,230 + 16,740 + 35,062 + 85,128 + 30,899 + 94,732 + 39,740 + 63,158 + 64,384 + 40,628 + 45,646 = 890,365)

ARR at risk in window (High + Medium):
- High: C-0BBE3E60 $30,993 + C-0F5D2323 $90,647 + C-0EC6999D $79,419 = $201,059
- Medium: C-0B20DB64 $21,770 + C-0FD551AB $48,815 + C-0F9F8F13 $46,230 = $116,815
- Total at risk: $201,059 + $116,815 = $317,874

Additional exposure outside the window: the three already-passed renewals (C-0B7D2C30, C-0BCDB8C2, C-0D2AB865) total $65,901 + $54,427 + $38,022 = $158,350, all showing High-risk usage patterns, but their renewal outcomes are not in the provided data. Across all 20 accounts, total ARR on the books is $1,048,715.

MISSING DATA: renewal outcomes for the three passed dates; any contract/negotiation status, health scores, or support signals beyond usage; no company names beyond the given aliases.
ticket-theme-synthesis0.80
Plain-English question: What are the real themes in this quarter's support tickets, ignoring the existing tags?
Prompt — what the model saw
Synthesize the quarter's support tickets into 4 to 7 themes. Do not trust existing tags; read the text. Per theme: name, count, share, distinct accounts, ARR affected, two ticket ids, one-line recommendation. Separate broad patterns from single-account noise and rank by ARR exposure, not volume.
cost $0.0184 · 189s · in 2,933 / out 9,071 tok
✗ theme_count_range✓ counts_verifiable✓ single_account_flagged✓ arr_ranked✓ ticket_ids

✓ Correct Answer (ground truth)

Plain-English answer: 80 tickets in 4-7 themes — points-not-posting leads (20 tickets, 25%, $31,100 ARR affected) — each with counts, distinct accounts, ARR, and a recommendation.
The answer in one line: total_tickets = 80; quarter_window = 2026-06-01 through 2026-08-31
total_tickets80
noteexisting tags are noisy/unreliable; classify by text
quarter_window2026-06-01 through 2026-08-31
ranking_by_arr_exposurehris_provisioning, redemption_failures, billing_dispute, points_not_posting, slack_sync
sample_ticket_ids80 items — e.g. IC-460004, IC-460016, IC-460001, IC-460041
single_account_noise_themebilling_dispute
themes{"points_not_posting": {"count": 20, "share_pct": 25.0, "distinct_accounts": 9, "arr_affected": "31100.00", "example_ticket_ids": ["IC-460001", "IC-460002"], "single_account": false}, "slack_sync": {"
graded checkstheme_count_range · counts_verifiable · single_account_flagged · arr_ranked · ticket_ids

Model's actual answer

show the model's full answer
Read all 80 ticket bodies (Jun 1 – Aug 29, 2026); tags ignored — e.g., IC-460055 is tagged "billing" but describes HRIS provisioning, and IC-460020 tagged "billing" describes missing points. Six themes; ranked by ARR exposure (sum of distinct accounts' ARR per theme), not ticket count.

THEME 1 — HRIS provisioning failures (broad, enterprise-only)
- Count: 12 of 80 (15.0%)
- Distinct accounts: 3 — C-0B2213A9 (36,000), C-0F6C0F34 (30,000), C-0DDFC9A7 (48,000)
- ARR affected: 36,000 + 30,000 + 48,000 = $114,000
- Tickets: IC-460060, IC-460064
- Recommendation: Escalate to engineering as a P1 provisioning-pipeline defect — it touches every enterprise account in the file.

THEME 2 — Redemption/checkout/gift-card failures (broad)
- Count: 18 of 80 (22.5%)
- Distinct accounts: 7 — C-0CEF69FD (8,900), C-0B827671 (10,700), C-0FCCD2DF (9,600), C-0F876796 (8,700), C-14264ABD (11,000), C-0B0F1BAB (10,300), C-0D9CA315 (9,600)
- ARR affected: 8,900+10,700+9,600+8,700+11,000+10,300+9,600 = $68,800
- Tickets: IC-460025, IC-460038
- Recommendation: Fix checkout timeout plus the points-deducted-without-delivery state; the latter is a trust issue requiring proactive refunds/credit.

THEME 3 — Invoice seat-count discrepancies (single-account noise, highest-value account)
- Count: 10 of 80 (12.5%)
- Distinct accounts: 1 — C-0E9C27D1
- ARR affected: $52,000
- Tickets: IC-460071, IC-460079
- Recommendation: Assign a named CSM/engineering owner to C-0E9C27D1; ten repeat invoices on the same 200-vs-150 seat error signals an unresolved billing-config bug.

THEME 4 — Renewal charged at wrong tier price (single-account noise, same account)
- Count: 6 of 80 (7.5%)
- Distinct accounts: 1 — C-0E9C27D1
- ARR affected: $52,000 (same account as Theme 3; not additive)
- Tickets: IC-460078, IC-460073
- Recommendation: Fold into the same C-0E9C27D1 war room — renewal pricing and seat count likely share one root cause in the billing config.

THEME 5 — Points not posting to balances (broad, mostly SMB)
- Count: 20 of 80 (25.0%) — highest volume, mid exposure
- Distinct accounts: 9 — C-0D3278C7 (3,500), C-0BF20542 (4,500), C-0D0B047C (4,500), C-0BE96399 (2,700), C-0D284E42 (3,400), C-0D6CC8E3 (4,200), C-21FEBCBB (2,900), C-0DD0626C (2,500), C-0B2895EF (2,900)
- ARR affected: 3,500+4,500+4,500+2,700+3,400+4,200+2,900+2,500+2,900 = $31,100
- Tickets: IC-460004, IC-460007
- Recommendation: Treat as one recognition-ledger delivery defect (individual and team-wide reports share the symptom); add balance-drift monitoring before points send.

THEME 6 — Slack integration breakage (broad, mid-market)
- Count: 14 of 80 (17.5%)
- Distinct accounts: 4 — C-0B843542 (4,400), C-10A56B0F (5,400), C-0BA71F12 (3,900), C-8C2E8F00 (5,200)
- ARR affected: 4,400+5,400+3,900+5,200 = $18,900
- Tickets: IC-460049, IC-460046
- Recommendation: Fix token refresh so re-auth sticks and slash commands stop erroring; one root cause plausibly explains all four symptom variants.

Broad patterns vs single-account noise:
- Broad (multi-account, systemic): Themes 1, 2, 5, 6 — 64 tickets (80.0%), 21 distinct accounts, $232,800 combined unique ARR.
- Single-account noise: Themes 3–4 — 16 tickets (20.0%) but all from C-0E9C27D1 ($52,000), the largest account in the file; high churn risk despite being one account.

Coverage check: 14+12+10+6+18+20 = 80 tickets; distinct-account ARR across themes = 114,000+68,800+52,000+31,100+18,900+52,000 − 52,000 (C-0E9C27D1 counted once) = $284,800, matching the file's total unique-account ARR.
marketing 5 tests
case-study-match0.75
Plain-English question: Which three existing customers are the best case-study matches for this prospect?
Prompt — what the model saw
Given the prospect, rank the three most similar existing customers with public case studies for use as social proof. Similarity on industry, size band, use case, region. Cite the fields that drove each match. Do not name any customer without a case study and never include billing data or contact names.
cost $0.0059 · 32s · in 0 / out 1,701 tok
✓ top3✗ decoys_excluded✓ cites_fields✓ no_billing

✓ Correct Answer (ground truth)

Plain-English answer: Three ranked matches on industry, size, use case, and region (e.g. C-11C31562), with the higher-scoring decoy explicitly excluded for a mismatched field.
The answer in one line: prospect_alias = C-82AF3719
constraintsno_billing_data, no_contact_names, only_customers_with_has_case_study_true_named
decoys_excludedcompany_alias=C-D6217CAA; similarity_score=4; reason=no_case_study | company_alias=C-C153868F; similarity_score=4; reason=no_case_study
prospect_aliasC-82AF3719
prospect_fields_usedindustry, size_band, use_case, region
top3company_alias=C-11C31562; similarity_score=3 | company_alias=C-64171065; similarity_score=3 | company_alias=C-A13C193D; similarity_score=3
graded checkstop3 · decoys_excluded · cites_fields · no_billing

Model's actual answer

show the model's full answer
Method: scored each case-study customer (has_case_study=true) against prospect C-82AF3719 (Technology, Mid-Market, employee_recognition, NA-West) on 4 equally weighted fields — industry, size_band, use_case, region — 1 point per match, max 4.

Case-study customers only (8 of 25 rows qualify: C-FEFCC020, C-11209339, C-9D69D147, C-11C31562, C-CD4829A7, C-0CE315B2, C-64171065, C-A13C193D).

Scores:
- C-64171065: industry ✓ + size_band ✓ + use_case ✓ + region ✗ = 3/4
- C-11C31562: industry ✗ + size_band ✓ + use_case ✓ + region ✓ = 3/4
- C-A13C193D: industry ✓ + size_band ✓ + use_case ✗ + region ✓ = 3/4
- C-9D69D147: ✗ + ✓ + ✗ + ✓ = 2/4
- C-CD4829A7: ✓ + ✓ + ✗ + ✗ = 2/4
- C-FEFCC020, C-11209339, C-0CE315B2: 0/4

Three-way tie at 3/4. Tiebreak: use_case match weighted highest (strongest social-proof signal for an employee_recognition buyer), then industry, then region — applied consistently.

RANKING

1. C-64171065 — 3/4
   Matched fields: industry (Technology), size_band (Mid-Market), use_case (employee_recognition). Differing field: region (NA-East vs prospect NA-West).

2. C-11C31562 — 3/4
   Matched fields: size_band (Mid-Market), use_case (employee_recognition), region (NA-West). Differing field: industry (Manufacturing vs Technology).

3. C-A13C193D — 3/4
   Matched fields: industry (Technology), size_band (Mid-Market), region (NA-West). Differing field: use_case (retention vs employee_recognition).

Notes and limitations:
- No customer in the dataset matches the prospect on all 4 fields; the closest available is 3/4. No Technology/Mid-Market/employee_recognition/NA-West customer exists in this file (the only exact-use-case NA-West rows are non-case-study: C-D6217CAA, C-C153868F, C-37C484FD, C-EB10E4AA — excluded per the case-study rule).
- Data missing: no revenue, headcount, or deal-size fields were provided, so size-band similarity is band-level only. No weighting scheme was specified; the equal-weight + use-case tiebreak above is my stated assumption.
- No billing data or contact names appear in the provided data or this output.
channel-attribution-mathPASS 1.00
Plain-English question: Compute cost per SQM, cost per SQO, and pipeline per dollar for every channel over six months.
Prompt — what the model saw
Compute channel performance for the trailing 6 months. Paid: spend, SQMs, SQOs, cost per SQM, cost per SQO, SQM-to-SQO rate, pipeline amount, pipeline per dollar. Organic: volume, SQO rate, pipeline. Report a channel with spend and zero SQMs as undefined, not zero. Flag rows where SQO date precedes SQM date. Finish with a reallocation recommendation and your confidence given sample sizes.
cost $0.0317 · 329s · in 1,950 / out 19,096 tok
✓ undefined_channel✓ undefined_not_zero✓ flagged_rows✓ best_channel✓ reallocation

✓ Correct Answer (ground truth)

Plain-English answer: LinkedIn Ads at $960/SQM and $3,000/SQO, organic search at $90,000 pipeline on zero spend, paid social flagged at $18,000 with zero SQMs, and a channel winner named.
The answer in one line: best_pipeline_per_dollar_channel = paid_search
best_pipeline_per_dollar_channelpaid_search
confidence_notewebinars and referral have small SQM samples (<20); treat rates as directional.
flagged_rows_sqo_before_sqmCT-000041, CT-000044
months2026-03, 2026-04, 2026-05, 2026-06, 2026-07, 2026-08
per_channel{"linkedin_ads": {"spend_usd": 24000, "sqms": 25, "sqos": 8, "pipeline_amount": 96000, "cost_per_sqm": 960.0, "cost_per_sqo": 3000.0, "sqm_to_sqo_rate": 0.32, "pipeline_per_dollar": 4.0}, "organic_sea
reallocation_recommendationfrom=paid_social; to=paid_search; rationale=paid_social has spend but zero SQMs (cost per SQM undefined, not zero); paid_search has the highest pipeline per dollar.
undefined_channelspaid_social
graded checksundefined_channel · undefined_not_zero · flagged_rows · best_channel · reallocation

Model's actual answer

show the model's full answer
CHANNEL PERFORMANCE — TRAILING 6 MONTHS (2026-03 through 2026-08)
All 122 contacts in the file have SQM dates within this window, so all are included. Spend file covers 2026-03 to 2026-08 for 4 channels; organic_search and referral have no spend rows (treated as no spend data, not zero spend).

PAID CHANNELS

paid_search
- Spend: 6 × $6,000 = $36,000
- SQMs: 40
- SQOs: 18
- Cost/SQM: $36,000 / 40 = $900
- Cost/SQO: $36,000 / 18 = $2,000
- SQM→SQO: 18 / 40 = 45.0%
- Pipeline: 18 × $40,000 = $720,000
- Pipeline per $: $720,000 / $36,000 = $20.00

linkedin_ads
- Spend: 6 × $4,000 = $24,000
- SQMs: 25
- SQOs: 8
- Cost/SQM: $24,000 / 25 = $960
- Cost/SQO: $24,000 / 8 = $3,000
- SQM→SQO: 8 / 25 = 32.0%
- Pipeline: 8 × $12,000 = $96,000
- Pipeline per $: $96,000 / $24,000 = $4.00
- Excluding the 2 flagged rows below: 6 SQOs, cost/SQO $24,000/6 = $4,000, rate 6/25 = 24.0%, pipeline $72,000, pipeline/$ $3.00

paid_social
- Spend: 6 × $3,000 = $18,000
- SQMs: 0 (no contacts attributed in first-touch file)
- Cost/SQM: undefined (zero SQMs)
- Cost/SQO: undefined
- SQM→SQO: undefined
- Pipeline: $0
- Pipeline per $: $0.00

webinars
- Spend: 6 × $1,500 = $9,000
- SQMs: 12
- SQOs: 5
- Cost/SQM: $9,000 / 12 = $750
- Cost/SQO: $9,000 / 5 = $1,800
- SQM→SQO: 5 / 12 = 41.7%
- Pipeline: 5 × $12,000 = $60,000
- Pipeline per $: $60,000 / $9,000 = $6.67

ORGANIC CHANNELS (no spend rows provided; pipeline-per-dollar not computable)

organic_search
- Volume: 30 SQMs
- SQO rate: 10 / 30 = 33.3%
- Pipeline: 10 × $9,000 = $90,000
- Spend: not provided

referral
- Volume: 15 SQMs
- SQO rate: 6 / 15 = 40.0%
- Pipeline: 6 × $8,000 = $48,000
- Spend: not provided

DATA QUALITY FLAGS — SQO date precedes SQM date
- CT-000041 (linkedin_ads): SQM 2026-06-14, SQO 2026-06-09 — SQO 5 days before touch
- CT-000044 (linkedin_ads): SQM 2026-07-23, SQO 2026-07-18 — SQO 5 days before touch
Both are counted in LinkedIn's headline numbers above; the ex-flagged view is shown for comparison. (CT-000007 paid_search has SQO same day as SQM — not flagged.)

REALLOCATION RECOMMENDATION
1. Shift paid_social's $18,000 out of paid social pending an attribution audit — zero SQMs on 6 months of spend is either a tracking failure or true zero return; either way the dollars are idle.
2. Shift a portion of linkedin_ads budget into paid_search. Paid search leads every efficiency metric ($2,000/SQO, 45% conversion, $20 pipeline per dollar) vs LinkedIn at $3,000–$4,000/SQO and $3–4 pipeline per dollar even before the two date anomalies.
3. Hold webinars flat — best cost/SQM ($750) and strong cost/SQO ($1,800), but the sample is too small to scale confidently.
4. Invest in enabling organic_search and referral (content, customer advocacy) — 40%+ SQO rates on unattributed/zero spend; they're the cheapest pipeline in the file.

CONFIDENCE
- paid_search: HIGH — n=40 SQMs, 18 SQOs; stable monthly flow.
- organic_search: MODERATE — n=30, 10 SQOs.
- linkedin_ads: MODERATE-LOW — n=25, 8 SQOs, and 2 of 8 SQOs (25%) are date anomalies; conclusions about LinkedIn are sensitive to whether those rows are valid.
- referral: LOW-MODERATE — n=15, 6 SQOs.
- webinars: LOW — n=12, 5 SQOs; a 5/12 rate carries a wide interval (roughly 17–69% at 95%); do not scale on this sample.
- paid_social: NONE — n=0; cannot distinguish attribution breakage from channel failure.

Overall: the paid_search superiority call is well supported; the LinkedIn-vs-webinars ordering is not, given sample sizes and the flagged rows.
competitor-battlecard-updatePASS 1.00
Plain-English question: Update the Rivally battlecard with sourced pricing and where each side wins.
Prompt — what the model saw
Update the battlecard for the competitor. Sections: one-line positioning, pricing with source and date (newer source wins, note the conflict), where they win, where we win, objections and responses, recent changes, our 12-month win/loss record against them. Cite a snippet id for every factual claim. Rep opinion on a call is not a fact about the competitor. Anything from the old card you cannot re-source gets marked unverified.
cost $0.0115 · 94s · in 311 / out 5,486 tok
✓ competitor_named✓ pricing✓ win_loss✓ unverified_marked✓ citations

✓ Correct Answer (ground truth)

Plain-English answer: Rivally at $7 per user per month (the August source wins over the stale $5, conflict noted), a 13-7 win/loss record, and two unverified old claims flagged.
The answer in one line: competitor = Rivally
citation_ids_available25 items — e.g. S01, S02, S03, S04
competitorRivally
pricingvalue_per_user_per_month=7; source_snippet=S17; source_date=2026-08-12
rep_opinion_snippet_ids_not_factsS09, S21
unverified_old_claimsRivally lacks a Slack integration., Rivally was acquired by WorkHuman in 2025.
win_loss_12mowins=13; losses=7
graded checkscompetitor_named · pricing · win_loss · unverified_marked · citations

Model's actual answer

show the model's full answer
# Battlecard: Rivally (updated 2026-09-24; source data through 2026-09-03)

**Positioning:** Points-based recognition for mid-market, now pushing upmarket into EU enterprise on data residency and multi-language support (S02, S04, S12, S15).

## Pricing
- Current list: Recognition Starter **$7/user/mo, annual billing required** — pricing page, 2026-08-12 (S17). Newer source wins; supersedes $5/user/mo seen on pricing pages 2026-01-20 (S03) and 2026-04-01 (S08).
- Conflict to note: S13 (2026-06-02, call_notes) records a **$6.50/user/mo quote to a 500-seat prospect** while list was still $5 (S08). The data does not explain the gap (higher tier vs. discount — unconfirmed).
- S18 (2026-08-14, call_notes): prospect-reported $7 list with a **15% discount offered on a 3-year term** — consistent with S17 on list price; the discount is deal-reported, not confirmed by any pricing page.
- Rivally Pulse: sold as a separate add-on, not bundled (S23, 2026-09-01); no add-on price appears in the data.

## Where they win
- EU: EU data residency GA + Dublin office (S15); pitched EU residency in a live deal (S05); strong with distributed EU teams, multi-language support praised (S12).
- Engagement: recognition feed repeatedly praised (S02, S16).
- Onboarding: setup under a week; Slack integration worked out of the box (S04).
- Support: response time under 4 hours praised (S22).

## Where we win
- Analytics depth: 800-seat prospect chose Bonusly over Rivally citing analytics depth (S25); Rivally analytics called limited (S02), dashboards basic vs. enterprise tools (S07), and exports are CSV-only, making migration off Rivally hard (S20).
- Enterprise admin: no SCIM provisioning; manual user management painful (S10); no bulk recognition editing (S24); admin tooling lags peers (S16).

## Objections & responses
- "Rivally has EU data residency." True as of 2026-07-01 (S15). Counter on enterprise-scale gaps: no SCIM (S10), weak analytics (S25).
- "Rivally is cheaper." List is now $7/user/mo (S17), up from $5 (S03, S08). A 15%/3-year discount appears in one prospect report only (S18, unconfirmed). Pivot to total value: analytics depth (S25) and admin automation (S10, S24).
- "Their feed is more engaging." Acknowledged (S02, S16). Counter with analytics (S07, S25) and admin tooling (S24).
- "Setup is faster with Rivally." Supported for them (S04). No data exists on our setup time — do not claim either way.

## Recent changes
- Price increase: Starter $5 → $7/user/mo (2026-08-12, S17).
- Rivally Pulse exits beta as a separately priced add-on (S06 2026-03-05 launch; S23 2026-09-01 GA).
- Microsoft Teams app v2 in public preview (S19, 2026-08-20).
- EU expansion: ex-Workday VP EMEA hired to lead Europe (S11, 2026-05-09); Dublin office + EU data residency GA (S15, 2026-07-01).
- Series C, $40M, led by Northgate Ventures (S01, 2025-11-04).

## Our 12-month win/loss record (2025-09 → 2026-08; all 20 recorded deals)
- **13 wins – 7 losses = 65% win rate** (13/20).
- Losses (7): Deal-7767F5 (2025-09), Deal-D263E0 (2025-11), Deal-935746 (2025-12), Deal-9066A6 (2026-03), Deal-5645A5 (2026-04), Deal-72A02F (2026-04), Deal-C6FFAA (2026-05).
- Wins (13): Deal-072E31 (2025-09), Deal-A9FD43 (2025-10), Deal-F65C8F (2025-10), Deal-7AA785 (2025-11), Deal-44C524 (2025-12), Deal-0D0CD6 (2026-01), Deal-E46EAB (2026-01), Deal-D5B790 (2026-02), Deal-1D2392 (2026-02), Deal-5C636E (2026-03), Deal-67BE14 (2026-06), Deal-1B6969 (2026-07), Deal-F03E7B (2026-08).
- Pattern: 0–3 stretch across 2026-04/05, then won every month 2026-06 through 2026-08 (3–0).
- Note: S25 (analytics-depth win) is not linked to any deal alias in the deals file — treat as color, not a record entry. No loss reasons are recorded in the data.

## Old-card claims — re-source status
- "Points-based recognition for mid-market" — re-sourced (S02, S04). Retained.
- "$5/user/mo, annual (as of 2026-01)" — re-sourced for that date (S03) but superseded by S17 ($7). Updated above.
- "Rivally lacks a Slack integration" — **contradicted**: S04 (2026-02-02) reports Slack integration worked out of the box. Removed.
- "Acquired by WorkHuman in 2025" — **UNVERIFIED**: no snippet supports it; S11 (ex-Workday VP hire) is not evidence of an acquisition. Marked unverified.
- "Strong in EU enterprise with multi-language support" — re-sourced (S12); reinforced by S15.

**Excluded as rep opinion (not competitor fact):** S09 (AE opinion on UI), S21 (AE opinion on aggressive discounting). Neither is used as a factual claim anywhere above.
nurture-sequence-diagnosisPASS 1.00
Plain-English question: What is broken in each email sequence — rates, overlaps, tracking errors?
Prompt — what the model saw
Review the sequences. Per sequence: sent, open rate, reply rate, meeting rate, weakest step. Identify tracking errors (opened above sent), audience overlap between sequences, and the failure mode for anything under 2 percent reply. One change per weak sequence and which to fix first. Under 350 words.
cost $0.0289 · 219s · in 10,545 / out 12,765 tok
✓ tracking_error✓ overlap_pair✓ weakest_step✓ fix_priority

✓ Correct Answer (ground truth)

Plain-English answer: Per-sequence open, reply, and meeting rates with the weakest step (all three weakest at step 3), an opened-above-sent tracking error, and audience overlap flagged.
The answer in one line: fix_priority_first = Cold Outbound - HR Leaders
all_pairwise_overlaps['Cold Outbound - HR Leaders', 'Cold Outbound - People Ops'], ['Expansion Nurture', 'New Logo Nurture']
audience_overlap_pairCold Outbound - HR Leaders, Cold Outbound - People Ops
failure_modeCold Outbound - HR Leaders=opens healthy (~37%) but replies near zero from step 1 -> body/CTA not landing, not a deliverability or subject-line problem
fix_priority_firstCold Outbound - HR Leaders
per_sequence{"New Logo Nurture": {"sent": 1386, "opened": 490, "replied": 90, "meetings": 27, "open_rate": 0.3535, "reply_rate": 0.0649, "meeting_rate": 0.0195, "weakest_step": 3}, "Expansion Nurture": {"sent": 8
tracking_errorsequence=Expansion Nurture; step=2; sent=300; opened=340; issue=opened exceeds sent
under_2pct_replyCold Outbound - HR Leaders
graded checkstracking_error · overlap_pair · weakest_step · fix_priority

Model's actual answer

show the model's full answer
**Per-sequence metrics** (rates = count ÷ sent, summed across steps)

- New Logo Nurture: sent 1,386 (500+458+428); open 35.4% (490/1,386); reply 6.5% (90/1,386); meeting 1.9% (27/1,386). Weakest step: 3 — 4.2% reply (18/428), 28.0% open (120/428).
- Expansion Nurture: sent 875 (300+300+275); open 64.6%* (565/875); reply 6.7% (59/875); meeting 1.4% (12/875). Weakest step: 3 — 4.4% reply (12/275), 34.5% open (95/275). *Open rate unreliable (see tracking error).
- Cold Outbound - HR Leaders: sent 1,785 (600+595+590); open 30.5% (545/1,785); reply 0.45% (8/1,785); meetings 0/1,785 = 0%. Weakest step: 3 — 0.17% reply (1/590), 0 meetings anywhere.
- Cold Outbound - People Ops: sent 1,163 (400+386+377); open 29.2% (340/1,163); reply 2.5% (29/1,163); meeting 0.5% (6/1,163). Weakest step: 3 — 1.6% reply (6/377); steps 1–2 are 3.5% (14/400) and 2.3% (9/386).

**Tracking error:** Expansion Nurture step 2 shows opened 340 > sent 300 (113%) — impossible; open counting is inflated (duplicate or multi-open attribution). Its 64.6% open rate is not trustworthy. Only row affected.

**Audience overlap:** 21 contacts enrolled in two sequences each.
- 19 in both "Cold Outbound - HR Leaders" and "Cold Outbound - People Ops": CT-000849, CT-000884, CT-000890, CT-000908, CT-001033, CT-001097, CT-001101, CT-001103, CT-001130, CT-001153, CT-001159, CT-001217, CT-001227, CT-001236, CT-001255, CT-001258, CT-001285, CT-001311, CT-001345.
- 2 in both "Expansion Nurture" and "New Logo Nurture": CT-000301, CT-000624.

**Failure mode (<2% reply):** Cold Outbound - HR Leaders. Opens are healthy (30.5%) but replies are 0.45% with zero meetings — deliverability is fine; the message content fails to convert opens into replies. Double-enrollment (19 contacts also in People Ops) compounds fatigue in the same audience.

**One change per weak sequence:**
- Cold Outbound - HR Leaders: rewrite the message body/CTA (content fix, not deliverability).
- Cold Outbound - People Ops: rewrite step 3 (1.6% reply, the only sub-2% step).

**Fix first:** Cold Outbound - HR Leaders — worst reply rate (0.45%), zero meetings, largest volume (1,785 sends), and a contaminated audience; dedupe against People Ops, then rewrite.
weekly-marketing-goals-rollupPASS 1.00
Plain-English question: Are we on pace on the quarter's marketing goals?
Prompt — what the model saw
Produce the weekly marketing goals update for the active quarter: SQMs, SQOs, DS2s, closed-lost MIA rate, same-quarter close count, active pipeline coverage against target. For each: QTD actual, target, delta, pace (on/behind/ahead given days elapsed). One paragraph on what moved this week.
cost $0.0098 · 87s · in 0 / out 4,477 tok
✓ metrics_present✓ deltas✓ coverage_ratio✓ pace✓ paragraph

✓ Correct Answer (ground truth)

Plain-English answer: SQMs ahead of pace, SQOs on pace, DS2s behind (35 short of expected at day 66 of 92), with deltas and same-quarter closes.
The answer in one line: days_in_quarter = 92; days_elapsed = 66; coverage_ratio_active_pipeline = 0.75; elapsed_fraction = 0.7174 (+1 more below)
days_in_quarter92
coverage_ratio_active_pipeline0.75
days_elapsed66
elapsed_fraction0.7174
metrics{"SQMs": {"qtd_actual": 230, "target": 300, "delta": -70, "expected_at_day_elapsed": 215.22, "pace": "ahead"}, "SQOs": {"qtd_actual": 84, "target": 120, "delta": -36, "expected_at_day_elapsed": 86.09,
pace_ruleexpected = target * days_elapsed/days_in_quarter; within +/-5% of expected = on, above = ahead, below = behind; lower_better metrics invert.
this_week_paragraph_required1
graded checksmetrics_present · deltas · coverage_ratio · pace · paragraph

Model's actual answer

show the model's full answer
WEEKLY MARKETING GOALS UPDATE — Q3-2026 (Jul 1–Sep 30, 2026)
Days elapsed: 66 of 92 = 66/92 = 0.7174 (71.7% of quarter). Pro-rated expected = target × 0.7174.

Metric                  | QTD actual | Target | Delta (act−tgt) | Expected @71.7% | Pace
------------------------|------------|--------|-----------------|-----------------|------
SQMs                    | 230        | 300    | −70             | 215.2           | AHEAD (+14.8; 76.7% attainment)
SQOs                    | 84         | 120    | −36             | 86.1            | BEHIND (−2.1; 70.0%)
DS2s                    | 40         | 75     | −35             | 53.8            | BEHIND (−13.8; 53.3%)
Closed-lost MIA rate    | 0.20 (5/25)| 0.10   | +0.10           | n/a (rate)      | BEHIND — 2× the 10% ceiling (lower_better)
Same-quarter closes      | 10         | 20     | −10             | 14.3            | BEHIND (−4.3; 50.0%)
Active pipeline vs tgt  | $3,000,000 | $4,000,000 | −$1,000,000 | $2,869,565      | AHEAD (+$130,435; 75.0%)

Arithmetic:
- Pace fraction: 66/92 = 0.7174.
- SQMs: 300 × 0.7174 = 215.2; 230 − 215.2 = +14.8 → ahead. Attainment 230/300 = 76.7%.
- SQOs: 120 × 0.7174 = 86.1; 84 − 86.1 = −2.1 → behind. Attainment 84/120 = 70.0%.
- DS2s: 75 × 0.7174 = 53.8; 40 − 53.8 = −13.8 → behind. Attainment 40/75 = 53.3%.
- MIA rate: 5/25 = 0.20; 0.20 − 0.10 = +0.10 above target. Lower is better, so off-track. No time-pro-rating applied to a rate. Arithmetically, with 5 MIA already booked, ending at ≤0.10 requires total closed-lost ≥ 50 with zero further MIA (5/50 = 0.10).
- Same-quarter closes: 20 × 0.7174 = 14.3; 10 − 14.3 = −4.3 → behind. Attainment 10/20 = 50.0%.
- Pipeline: $4.0M × 0.7174 = $2.87M; $3.0M − $2.87M = +$0.13M → ahead. Attainment 3.0/4.0 = 75.0%.
- Required run-rate over remaining 26 days (92 − 66): 70 SQMs (2.7/day), 36 SQOs (1.4/day), 35 DS2s (1.3/day), 10 same-quarter closes (0.4/day), $1.0M pipeline ($38.5k/day).

What moved this week: The data provided is QTD cumulative only — no prior-week snapshot or week-over-week deltas — so what specifically moved this week cannot be determined from the inputs, and I won't speculate. What the QTD position shows: active pipeline (75.0%) and SQMs (76.7%) are the two metrics running ahead of the 71.7% elapsed pace; SQOs (70.0%) are essentially on pace (−2.1 vs pro-rated); the laggards are DS2s (53.3%) and same-quarter closes (50.0%), both materially behind with 26 days left; and the closed-lost MIA rate sits at 20% (5 of 25 losses), double the 10% target, recoverable only if remaining losses are entirely non-MIA and total losses reach ≥50. No deal or company aliases appear in the provided data, so none are cited.
communication 5 tests
ceo-slack-compressionPASS 1.00
Plain-English question: Compress the forecast analysis into one Slack message the CEO can read in ten seconds.
Prompt — what the model saw
Compress the forecast analysis below into a single Slack message for the CEO. Rules: at most three sentences; under 80 words; lead with the decision or the so-what; keep every distinct number that matters and the one caveat that changes the conclusion; cut any sentence that does not carry a claim, a number, or a directive. No em dashes, no throat-clearing phrases.
cost $0.0130 · 158s · in 0 / out 6,727 tok
✓ max_3_sentences✓ under_80_words✓ weighted_present✓ caveat_present✓ no_em_dash

✓ Correct Answer (ground truth)

Plain-English answer: Three sentences, under 80 words, decision first, every distinct number kept plus the one caveat — all checked by deterministic text rules.
The answer in one line: weighted_number = 115976; must_keep = data-quality caveat — deleting the bad news fails
weighted_number115,976
limitmax 3 sentences, max 80 words, no em-dashes
must_keepdata-quality caveat — deleting the bad news fails
graded checksmax_3_sentences · under_80_words · weighted_present · caveat_present · no_em_dash

Model's actual answer

show the model's full answer
Q3 forecast of 115,976.75 (44,729 COMMIT + 35% of 203,565 BEST_CASE) looks precise and is not. All 7 COMMIT deals have empty why-buys, so the committed number has no documented reason to buy behind it. Owner is blank on 85 of 86 deals, so no rep rollup or quota check; 32 of 86 worth 227,575 close in October, including COMMIT Deal-D348E1 at 13,770.
stalled-deal-followup-emailPASS 1.00
Plain-English question: Write a follow-up email for this stalled deal — under 120 words, one prior touch, one ask.
Prompt — what the model saw
Write a follow-up email for this stalled deal. Under 120 words. Reference one specific prior touch from the engagement history, make one clear ask, and match the tone of the prior thread. No invented details.
cost $0.0046 · 19s · in 0 / out 795 tok
✓ word_cap✓ references_touch✓ single_ask

✓ Correct Answer (ground truth)

Plain-English answer: An email that references the July 28 milestone-awards demo, makes one clear ask, and invents nothing — Deal-0D2F7A, 31 days since last engagement.
The answer in one line: amount = 5100; word_cap = 120; days_since_last_engagement = 31; deal_alias = Deal-0D2F7A (+4 more below)
amount5,100
days_since_last_engagement31
word_cap120
allowed_facts14 items — e.g. Deal-0D2F7A, $5,100, DS3, Alex Franklin
allowed_newENG-00001, ENG-00002, ENG-00003, ENG-00004, ENG-00005
deal_aliasDeal-0D2F7A
must_referenceengagement_id=ENG-00002; date=2026-07-28; type=meeting; why=most substantive prior touch; concrete prospect reactions to milestone awards and points catalog
ownerAlex Franklin
single_cta1
stageDS3
tonewarm, concise, professional (matches prior thread)
graded checksword_cap · references_touch · single_ask

Model's actual answer

show the model's full answer
Subject: Following up — pricing for 150 seats

Hi [Name],

Hope things are going well. Following up on my August 5 email, which included pricing for 150 seats after the July 28 demo with your People team — I know priorities shift, so no worries if it slipped.

One ask: would a quick 20-minute call in the next week or two work to walk through the pricing and answer any remaining questions from the demo?

Happy to send anything else that would help in the meantime.

Best,
Alex Franklin

(93 words. References the Aug 5 pricing email / July 28 demo touch; single ask is a short call; recipient name not in the data, so a placeholder is used.)
funnel-digest-paragraph0.33
Plain-English question: Write the weekly GTM digest — one paragraph per team, every number traceable.
Prompt — what the model saw
Write the weekly GTM digest: one paragraph per team (Marketing, Sales, CS, PLG) in Bonusly brand voice. Every number must trace to the extracts. No customer contact names, no employee performance commentary, no compensation. Four paragraphs, under 400 words total.
cost $0.0044 · 15s · in 0 / out 624 tok
✗ four_paragraphs✓ word_cap✗ teams

✓ Correct Answer (ground truth)

Plain-English answer: Four paragraphs (Marketing, Sales, CS, PLG) using only the allowed numbers, no contact names, no performance commentary, under 400 words.
The answer in one line: word_cap = 400
word_cap400
allowed_numbers16 items — e.g. 2, 3, 9, 08
bannedcustomer contact names, employee performance commentary, compensation
required_paragraphsMarketing, Sales, CS, PLG
graded checksfour_paragraphs · word_cap · teams

Model's actual answer

show the model's full answer
I'll load the Bonusly brand skill first since this is Bonusly-branded content.
partner-digest-sectionPASS 1.00
Plain-English question: Write the partner digest so every partner in the roster appears exactly once.
Prompt — what the model saw
Write the partner digest for the period. Every partner in the roster appears once: active plays get status with numbers, quiet partners get one line. Partner-sourced pipeline count and amount must match the deal data. Nothing omitted, nothing added.
cost $0.0061 · 39s · in 0 / out 1,844 tok
✓ all_partners✓ pipeline_amount✓ quiet_one_line

✓ Correct Answer (ground truth)

Plain-English answer: Four active partners with status and numbers, four quiet partners with one line each, and partner-sourced pipeline that reconciles to the deal data.
The answer in one line: rule = every roster partner named exactly once; quiet partners get one line; no partners outside roster.
active_no_dealsWorkWell Group
partner_sourced_pipelinecount=5; amount=550000
per_partner_pipeline{"Apex Rewards Co": {"count": 2, "amount": 275000}, "HRCloud Partners": {"count": 1, "amount": 140000}, "CultureBridge": {"count": 2, "amount": 135000}}
quiet_partnersRecogniq, KudosWave, PeopleFirst Advisors, TotalPerk
rosterpartner_name=Apex Rewards Co; status=active | partner_name=HRCloud Partners; status=active | partner_name=CultureBridge; status=active | partner_name=WorkWell Group; status=active | partner_name=Recogniq; status=quiet | partner_name=KudosWave; status=quiet … (+2 more)
ruleevery roster partner named exactly once; quiet partners get one line; no partners outside roster.
graded checksall_partners · pipeline_amount · quiet_one_line

Model's actual answer

show the model's full answer
Partner Digest — 2026-08-24 to 2026-09-04

PARTNER-SOURCED PIPELINE: 5 deals, $550,000
Arithmetic: 180,000 + 95,000 + 140,000 + 60,000 + 75,000 = 550,000

ACTIVE PARTNERS

Apex Rewards Co — 2 sourced deals, $275,000 (Deal-DDAAF2 $180,000 + Deal-2CF33E $95,000). Co-webinar locked for 09-15; both opps logged with UTM Source = Partner, both in DS1 (Slack 2026-08-25).

HRCloud Partners — 1 sourced deal, $140,000 (Deal-F1CDA5). Integration referral closed security review; opp moved to DS2 (Slack 2026-08-27).

CultureBridge — 2 sourced deals, $135,000 (Deal-096E1D $60,000 + Deal-067213 $75,000). Lunch-and-learn produced two sourced opps, both early stage (Slack 2026-08-29).

WorkWell Group — 0 sourced deals. Wants to restart the joint playbook in Q4; planning call booked for 09-09 (Slack 2026-09-02). No deals in partner_deals.csv, consistent with the Slack note.

QUIET PARTNERS

Recogniq — no sourced deals or activity reported this period.
KudosWave — no sourced deals or activity reported this period.
PeopleFirst Advisors — no sourced deals or activity reported this period.
TotalPerk — no sourced deals or activity reported this period.

Cross-check: Slack reports 2 (Apex) + 1 (HRCloud) + 2 (CultureBridge) + 0 (WorkWell) = 5 sourced opps, matching the 5 deals in partner_deals.csv. All 5 deals carry utm_source = Partner.
rfp-answer-from-policyPASS 1.00
Plain-English question: Answer the 10 RFP items from the policy excerpts only — no guessing.
Prompt — what the model saw
Answer the 10 questionnaire items from the policy excerpts only. Cite the excerpt id for every answer. Where the excerpts do not answer the question, mark it unanswerable and say what document would; do not guess. Tone: vendor response, direct.
cost $0.0047 · 22s · in 0 / out 868 tok
✓ answered_count✓ unanswerable_flagged✓ citations✓ no_guess

✓ Correct Answer (ground truth)

Plain-English answer: Seven answered with excerpt citations (Q1-Q7), three flagged unanswerable with the document that would answer them (Q8-Q10).
The answer in one line: answered_count = 7; rule = unanswerable items must be flagged and the document that would answer them named; do not guess.
answered_count7
answer_map{"Q1": ["P1"], "Q2": ["P1"], "Q3": ["P2"], "Q4": ["P3"], "Q5": ["P4"], "Q6": ["P5"], "Q7": ["P6"]}
ruleunanswerable items must be flagged and the document that would answer them named; do not guess.
unanswerableQ8, Q9, Q10
valid_citation_idsP1, P2, P3, P4, P5, P6, P7, P8
graded checksanswered_count · unanswerable_flagged · citations · no_guess

Model's actual answer

show the model's full answer
Vendor responses below. All answers sourced solely from the provided excerpts; excerpt ids cited per item.

Q1. Is customer data encrypted at rest? If so, how?
Yes. All customer data is encrypted at rest using AES-256. [P1]

Q2. Is data encrypted in transit?
Yes. Data in transit is encrypted with TLS 1.2 or higher. [P1]

Q3. Do you hold a SOC 2 report, and can we obtain a copy?
Yes. We maintain a SOC 2 Type II report, renewed annually. The report is available to prospects under NDA. [P2]

Q4. How long is customer data retained after contract termination?
Customer data is retained for the duration of the contract plus 90 days, after which it is deleted from production systems. [P3]

Q5. Where is your subprocessor list published and how are changes notified?
A current list of subprocessors is maintained at trust.bonusly.com. Customers receive 30 days notice before a new subprocessor is added. [P4]

Q6. Do you support SAML single sign-on?
Yes. SAML 2.0 single sign-on is available on the Enterprise plan (SCIM user provisioning is also available on that plan). [P5]

Q7. What is your production backup schedule and retention?
Production databases are backed up daily. Backups are encrypted and retained for 35 days. [P6]

Q8. Do you hold an ISO 27001 certification?
Unanswerable from the provided excerpts. The excerpts cover SOC 2 [P2] but make no statement regarding ISO 27001. This would be answered by the certifications/compliance section of the trust portal (trust.bonusly.com) or the company's formal compliance certifications documentation.

Q9. What is your contractual uptime SLA percentage?
Unanswerable from the provided excerpts. No excerpt states a contractual uptime SLA. This would be answered by the Master Services Agreement or the SLA exhibit/schedule to the contract.

Q10. Will you sign a HIPAA Business Associate Agreement?
Unanswerable from the provided excerpts. No excerpt addresses HIPAA or BAA willingness. This would be answered by the company's legal/contract terms — e.g., the MSA, a security/privacy addendum, or the trust portal's compliance documentation.
ops-maintenance 5 tests
skill-file-reconciliation0.86
Plain-English question: Audit a skill set against its manifest and flag every conflict with a severity and an action.
Prompt — what the model saw
Reconcile the skill set below against its manifest. Find and report (with severity CRITICAL/WARNING/INFO and an action type TRIM_DESC, UPDATE_BODY, DELETE_SKILL, MERGE, or REVIEW): (1) any two skills whose ALWAYS-trigger phrases overlap/duplicate, (2) any circular delegation chain between skills (name it), (3) any delegation target that does not exist (dangling, name it), (4) any version conflict and which skill should survive, (5) how many manifest descriptions exceed 1,024 characters, (6) any hardcoded page ids, dates, or person names in skill bodies, (7) manifest drift in BOTH directions (files with no manifest row, and manifest rows with no file). Make one proposal per finding; rewrite nothing.
cost $0.0829 · 417s · in 64,621 / out 24,332 tok
✗ duplicate_trigger_found✓ circular_chain✓ dangling_target✓ oversized_desc_count✓ hardcoded_values✓ manifest_drift✓ severity_and_action

✓ Correct Answer (ground truth)

Plain-English answer: A severity-ranked list of TRIM_DESC / UPDATE_BODY / DELETE_SKILL / MERGE / REVIEW actions covering every rule violation found.
The answer in one line: files_audited = 14; rank = CRITICAL / WARNING / INFO with one fix action per issue
files_audited14
findingsduplicate trigger, circular delegation chain, dangling file reference, oversized description, hardcoded values, manifest drift
rankCRITICAL / WARNING / INFO with one fix action per issue
graded checksduplicate_trigger_found · circular_chain · dangling_target · oversized_desc_count · hardcoded_values · manifest_drift · severity_and_action

Model's actual answer

show the model's full answer
RECONCILIATION REPORT — SignalForge skill set vs skill_manifest.csv
Scope: 14 manifest rows, 14 SKILL.md files provided. Aliases cited exactly as given. Nothing rewritten; one proposal per finding.

────────────────────────────────────────
1. ALWAYS-TRIGGER OVERLAPS / DUPLICATES
────────────────────────────────────────

F1 — CRITICAL / MERGE — comms-drafter ↔ email-drafter
Five literal phrase collisions plus one prefix collision:
  comms-drafter: "write me an email," "draft a follow-up," "help me reply," "what should I say," "bump email," "contract nudge"
  email-drafter: "write me an email," "draft a follow-up," "help me reply to this," "what should I say," "bump email," "contract nudge"
Both also claim pasted-message review/rewrite/rating requests. Zero disambiguating trigger language exists between them. Both bodies even carry the identical contract-follow-up benchmark email verbatim — evidence of common origin.
Proposal: MERGE into email-drafter (it survives: deal-strategy-coach hardwires delegation to it and it owns the Gmail-signature machinery); absorb comms-drafter's unique scopes (Intercom/support replies, partner/channel comms) as sections; delete comms-drafter.

F2 — CRITICAL / UPDATE_BODY — pipeline-intelligence-report ↔ weekly-pipeline-report
Collisions: "pipeline update" (weekly claims "run the pipeline update" verbatim; pipeline-intelligence-report claims "pipeline update"), "pipeline report" ("run the pipeline report" vs "generate the pipeline report"/"do the pipeline report"), "what's the pipeline look like" vs "what does pipeline look like". Both add "or any variation" catch-alls.next-to-close disambiguates itself against pipeline-intelligence-report, but nothing disambiguates these two against each other.
Proposal: UPDATE_BODY both ALWAYS blocks to be mutually exclusive (weekly-pipeline-report keeps weekly/summary/SQM-SQO-DS2/bookings-MTD anchors; pipeline-intelligence-report keeps "score the pipeline"/"full pipeline"/"pipeline intelligence"/"pipeline review") plus one explicit tie-break rule for the bare "pipeline report/update" phrases. (MERGE of the two skills is the alternative; they serve different consumers, so trigger rewrite is preferred.)

F3 — WARNING / REVIEW — next-to-close ↔ sales-forecast
next-to-close: "which deals are most likely to close," "what's closing this week," "closest to signature." sales-forecast: "what do we think we're going to close," "what's our number." Both claim the close-outlook ask. next-to-close routes overflow to pipeline-intelligence-report but never mentions sales-forecast; sales-forecast mentions neither.
Proposal: REVIEW — add mutual disambiguation (deal-level shortlist → next-to-close; quarter-level number → sales-forecast).

F4 — INFO / REVIEW — universal ALWAYS stacking
model-selection claims "ALWAYS run this skill at the start of every task, without exception"; analysis-validator claims "Never skip — even on quick check requests"; signalforge-feedback claims "ALWAYS trigger… Never skip." Three unconditional mandates, plus signalforge-claim-compressor claiming the same output universe as signalforge-feedback (sequential there, codified, so benign). Separately, no skill ever invokes model-selection — its Integration Protocol is dead wiring.
Proposal: REVIEW — confirm the intended stack order (model-selection → analysis-validator → claim-compressor → feedback) in one place, and either wire model-selection into the orchestrator or drop its universal claim.

────────────────────────────────────────
2. CIRCULAR DELEGATION CHAIN
────────────────────────────────────────

F5 — WARNING / REVIEW — deal-strategy-coach ↔ email-drafter (2-node cycle)
  deal-strategy-coach → email-drafter: "When drafting manager-to-prospect emails, use the email-drafter skill…"
  email-drafter → deal-strategy-coach: "For deal strategy, diagnosis, or coaching (not email drafting), use deal-strategy-coach instead."
comms-drafter feeds the cycle upstream (comms-drafter → deal-strategy-coach → email-drafter → deal-strategy-coach → …). Lane markers exist on both sides but do not structurally break the loop.
Proposal: REVIEW — make one direction terminal (e.g., email-drafter answers strategy questions inline or defers with a no-re-delegation note), or have deal-strategy-coach inline the signature logic it delegates for.

No other cycle found: pipeline-intelligence-report → closed-lost-analysis is one-way (closed-lost-analysis Mode 4 only notes it is "called from pipeline-intelligence-report"); next-to-close → pipeline-intelligence-report is one-way.

────────────────────────────────────────
3. DANGLING DELEGATION TARGETS (not in manifest)
────────────────────────────────────────
The manifest is the only inventory provided; any invoked skill absent from it is dangling below.

F6 — WARNING / REVIEW — bonusly-brand — invoked as a mandatory pre-step by comms-drafter ("Before drafting any communication, apply the bonusly-brand skill"), email-drafter, sales-forecast, and signalforge-claim-compressor. No manifest row.
Proposal: REVIEW — add to manifest or confirm it lives outside this skill set.

F7 — WARNING / REVIEW — prospect-research-multithreading — invoked by comms-drafter, email-drafter, and deal-strategy-coach (mandatory "Cross-skill handoff" section: "Invoke prospect-research-multithreading whenever…"). No manifest row.
Proposal: REVIEW — same as F6.

F8 — WARNING / REVIEW — skill-orchestrator — referenced by analysis-validator §11 (Three-Way Sync cascade) and signalforge-feedback (activation checklist: "Skill registered in skill-orchestrator"). No manifest row. Also unresolved: three artifacts named in analysis-validator §11 — CUSTOMER_DATA_REFERENCE, HUBSPOT_CONNECTOR_REFERENCE, SIGNALFORGE_PRODUCT_INSIGHT_SKILL.
Proposal: REVIEW — confirm existence; register or remove the references.

F9 — WARNING / REVIEW — the eight specialist skills in analysis-validator §12.4's delegation table, none in the manifest: bonusly-data-questions, bonusly-product-questions, bonusly-business-reporting-questions, bonusly-rewards-questions, bonusly-ppp-questions, bonusly-feature-flag-questions, bonusly-deal-desk-questions, bonusly-datadog-questions.
Proposal: REVIEW — the validator's "always delegate" path fails at runtime if these don't exist; add to manifest or prune the table.

F10 — INFO / REVIEW — signalforge-reports (org skill at /mnt/skills/organization/signalforge-reports/) — mandatory pre-build dependency for pipeline-intelligence-report and weekly-pipeline-report. Not in manifest (path-based; may live outside this set).
Proposal: REVIEW — confirm the path resolves; note it in the manifest if it is in scope.

F11 — INFO / REVIEW — write_to_snowflake MCP tool (signalforge-feedback Step 5) — self-declared undeployed ("Activation blocked on write_to_snowflake MCP tool deployment"). Known-dangling by design.
Proposal: REVIEW — track activation; until then Step 5 is dead code. (Tool-level deps such as Aligned:get_room_brief, ZoomInfo:account_research, BonuslyGPT:query_snowflake cannot be verified from the provided data.)

────────────────────────────────────────
4. VERSION CONFLICTS (and what survives)
────────────────────────────────────────

F12 — WARNING / UPDATE_BODY — analysis-validator v3.2 vs v3.6 (internal)
Frontmatter: "Version: 3.6," "Last Updated: May 9, 2026"; changelog tops at 3.6; pipeline-intelligence-report's footer independently cites "Analysis Validator v3.6." But the §7 trail template still prints "Validator: analysis-validator v3.2."
Survivor: v3.6 (three independent references vs one stale string).
Proposal: UPDATE_BODY — correct the §7 trail-template version string to v3.6.

F13 — CRITICAL / UPDATE_BODY — Gong transcript column: SNIPPET vs TRANSCRIPT
closed-lost-analysis Source 3 SQL selects "t.SNIPPET AS transcript_content" from GONG_TRANSCRIPTS_AGG. stale-pipeline-report states that table "has exactly two columns: CONVERSATION_KEY and TRANSCRIPT. There is no CALL_SPOTLIGHT_BRIEF, no SNIPPET… they do not exist and will error." analysis-validator §13.7 canonical routing uses TRANSCRIPT only.
Survivor: the TRANSCRIPT-only pattern (two skills + the validator agree). closed-lost-analysis's query errors at runtime as written.
Proposal: UPDATE_BODY — rewrite closed-lost-analysis's Source 3 SQL to select t.TRANSCRIPT.

F14 — WARNING / UPDATE_BODY — AE roster: Core 6 vs 5
analysis-validator §12.3 (Updated May 4, 2026): Core 6 AEs — Bryce Harmon 119337721, Hugo Lindqvist 77260721, Dana Mercer 83155923, Alex Franklin 84342457, Cole Ingram 83155924, Gavin Porter 1520255671 — with the rule that any "Core 6" filter must include all six IDs. pipeline-intelligence-report ("AE owner IDs (verified May 2026)") lists only 5 — Hugo Lindqvist is absent, with no note of departure.
Survivor: analysis-validator §12.3 (dated, doctrine-designated canonical roster; G2-F points ID resolution to it).
Proposal: UPDATE_BODY — pipeline-intelligence-report adopts/references the §12.3 roster or documents why Hugo Lindqvist is excluded.

F15 — WARNING / REVIEW — transcript-source doctrine contradiction
deal-strategy-coach and email-drafter prescribe HubSpot first → Granola second → Gong third. analysis-validator G1-D prescribes GONG_TRANSCRIPTS_AGG × GONG_HUBSPOT_MAP as "The only approved transcript source… No fallback. No substitute."
Proposal: REVIEW — scope the two rules explicitly (interactive coaching pull vs published-analysis validation) so they cannot both be "the" rule; as written they directly contradict.

────────────────────────────────────────
5. DESCRIPTIONS EXCEEDING 1,024 CHARACTERS
────────────────────────────────────────
Answer: 0 of 14.
Arithmetic from manifest description_chars: max value = 1006 (pipeline-intelligence-report) and 1006 (signalforge-claim-compressor); next = 1004 (partner-digest). 1006 < 1024, 1004 < 1024; all remaining 11 rows range 656–996. No row exceeds 1,024. No TRIM_DESC action is required on current numbers.

────────────────────────────────────────
6. HARDCODED PAGE IDs, DATES, PERSON NAMES IN BODIES
────────────────────────────────────────

F16 — CRITICAL / UPDATE_BODY — weekly-pipeline-report: hardcoded quarter and dates
Body bakes "Q2 (April 1 – June 30, 2026; total ≈ 64–65)" into the business-day step and "Q2 QTD vs Q1 full quarter" into funnel comparisons, while the skill is written as quarter-agnostic ("current month," "current quarter"). Run after June 2026 it paces against a closed quarter. (Changelog precedent for exactly this bug class exists in sales-forecast v1.0→v1.1.)
Proposal: UPDATE_BODY — compute the quarter window at runtime; delete the fixed Q2 anchor.

F17 — WARNING / UPDATE_BODY — weekly-pipeline-report: hardcoded person
"Ben Lavin · Demand Generation" in the H1 and "for Ben's review" / "deliver the HTML file to Ben" in the body — a single-user skill presented as a team skill.
Proposal: UPDATE_BODY — generalize to role-based language or explicitly declare personal scope in the description.

F18 — INFO / UPDATE_BODY — weekly-pipeline-report: static Q1 2026 financials embedded ($365,152 vs $475,000 plan, 77%; $2,490,532 vs $3,288,000 forecast, 76%) despite the skill reading the bookings spreadsheet live every run.
Proposal: UPDATE_BODY — drop the baked figures; pull Q1 context from the spreadsheet at runtime.

F19 — WARNING / UPDATE_BODY — closed-lost-analysis: hardcoded dated stats and anecdotes
"the 30-deal AI-field sample from May 2026: 10 of 10 deals"; interventions table percentages ("17% of losses had explicit hold/pause," "8% stated budget as reason; 14%+"); dated loss cases (Softheon May 2026, MinIO May 4–12, Estee Lauder, Aurora Innovation, GCash, LIFTOFF/Nestlé, Ozinga, Ethos Cannabis, StickerYou). These will read as current findings as data drifts.
Proposal: UPDATE_BODY — label the sample stats and percentages as historical exemplars with as-of dates, or recompute them live per run.

F20 — WARNING / REVIEW — analysis-validator: hardcoded people and anchors
§12.3 embeds a 19-name GTM roster with owner IDs (incl. Amani Phipps 210200121, John Thomas 78303262, Yasmin Wahid 89062643); G1-K and §10 hardcode "Manish or Amani" as the Finance escalation path; G1-J embeds anchors "~452,000 / ~110,097" (self-mitigated — the skill demotes anchors to calibration and re-queries live). Roster hardcoding directly conflicts with stale-pipeline-report's doctrine: "Never hardcode rep names or owner IDs."
Proposal: REVIEW — keep §12.3 as the single canonical roster other skills reference (never copy), and replace "Manish or Amani" with a role.

F21 — WARNING / UPDATE_BODY — partner-digest: hardcoded IDs and personalization
Confluence cloud ID 73fe98de-a4a3-4869-9f8a-bb1eeed4cf7f, space ID 1958248479, folder ID 2286616609, and six page IDs (2286321666, 2265382925, 2236940297, 2237825028, 2239365136, 2238283777); owner personalization throughout ("Amani's threads," Slack ID U03QLMBL7AR, "Owner: Amani Phipps"); five first-name contacts (Kelli, Jen Lee, Hani, Bryce, Sara).
Proposal: UPDATE_BODY — move IDs to a references file, generalize the owner, and mark contact names as runtime-resolved.

F22 — INFO / REVIEW — sales-forecast: hardcoded space ID 2232811524, parent page 2232582148, cloud ID, and "Alaina" in the Manager Forecast view (changelog shows v1.0 had "Elena hardcoded" — recurring bug class). Proposal: REVIEW — acceptable as publish constants; confirm the escalations are current.

F23 — INFO / REVIEW — signalforge-feedback: hardcoded page IDs 2295136266 (feedback log), 2234417154 (parent), 2247295002 (Build Log), spaceId 2232811524, cloudId. Proposal: REVIEW — verify pages still exist before first write.

F24 — INFO / REVIEW — deal-strategy-coach: hardcoded Confluence playbook page 2257879045 (AE Excellence Playbook April 2026) and routing names "Perseus"/"Farid" in the ICP section. Proposal: REVIEW — confirm the page ID and routing owners are current.

F25 — INFO / REVIEW — model-selection: registry last_checked 2026-05-19 with a 14-day self-update rule; the date itself is hardcoded by design. Proposal: REVIEW — confirm the self-update procedure actually executes; the registry is already past any 14-day window relative to the other skills' May–June 2026 timestamps.

Accepted constants (not findings): HubSpot deal stage IDs 150582536/150582537/150582538/150582539/1175632767 and org ID 1973303 in analysis-validator, next-to-close, pipeline-intelligence-report, stale-pipeline-report — stable system constants required for URLs/filters (pipeline-intelligence-report itself says "verify at run time — do not hardcode," which it then does; noted, tolerable).

────────────────────────────────────────
7. MANIFEST DRIFT (BOTH DIRECTIONS)
────────────────────────────────────────
Files with no manifest row: 0.
Manifest rows with no file: 0.
Arithmetic: 14 SKILL.md files provided ↔ 14 manifest rows; file basenames match declared_name 14/14; all declared_name values match each file's frontmatter name 14/14; no duplicates on either side. No drift in either direction.

────────────────────────────────────────
SUMMARY
────────────────────────────────────────
CRITICAL: F1 (comms-drafter/email-drafter trigger duplication — MERGE), F2 (pipeline-intelligence-report/weekly-pipeline-report trigger collision — UPDATE_BODY), F13 (closed-lost-analysis SNIPPET query will error — UPDATE_BODY), F16 (weekly-pipeline-report hardcoded Q2 2026 — UPDATE_BODY).
WARNING: F3, F5, F6–F9, F12, F14, F15, F17, F19, F20, F21.
INFO: F4, F10, F11, F18, F22–F25.
Count checks: descriptions >1,024 chars = 0; manifest drift = 0 both directions.
runbook-from-slack-threadPASS 1.00
Plain-English question: Turn the resolved incident thread into a runbook.
Prompt — what the model saw
Turn the resolved incident thread into a runbook: ordered steps with the exact command or action taken, who ran it, how success was verified, and rollback for each step that changed state. Every step must trace to a message in the thread. Add nothing the thread does not contain; where the thread is unclear, mark the step as needing confirmation.
cost $0.0076 · 53s · in 0 / out 2,930 tok
✓ ordered_steps✓ rollback✓ needs_confirmation✓ sourced

✓ Correct Answer (ground truth)

Plain-English answer: Ordered steps from the thread, each with the exact command, who ran it, how success was verified, and rollback where state changed.
The answer in one line: rule = no steps beyond what the thread contains; every step traces to a message id.
ambiguous_steporder=4; msg=M05; mark=needs-confirmation
ruleno steps beyond what the thread contains; every step traces to a message id.
stepsorder=1; msg=M01; actor=Bryce Harmon; action=acknowledge PagerDuty alert, take IC; state_change=False; verified_by=M01 | order=2; msg=M02; actor=Farid Osman; action=check queue depth; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M02 (48,213 pending) | order=3; msg=M04; actor=Farid Osman; action=pause auto-enqueue via feature flag; command=bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'; state_change=True; rollback=bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'; verified_by=M07 (queue depth falling) | order=4; msg=M05; actor=Elena Sinclair; action=cleared the dead set; state_change=True; needs_confirmation=thread gives no command or method for clearing the dead set | order=5; msg=M06; actor=Bryce Harmon; action=scale reward-worker to 6 replicas; command=kubectl scale deployment/reward-worker --replicas=6; state_change=True; rollback=kubectl scale deployment/reward-worker --replicas=3; verified_by=M07/M08 | order=6; msg=M08; actor=Cole Ingram; action=verify queue drained and error rate baseline; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M08 (returns 0) … (+2 more)
graded checksordered_steps · rollback · needs_confirmation · sourced

Model's actual answer

show the model's full answer
RUNBOOK — reward-worker queue backlog incident (source: #incident-2026-08-29-reward-queue, resolved 2026-08-29)

Context from thread: PagerDuty fired for reward-worker queue depth > 10k [M01]. Measured depth was 48,213 pending jobs vs. normal under 500 [M02]. Dead set held 112 jobs, all Redis::TimeoutError from ~13:58 [M03]. No root-cause fix appears in the thread; the runbook below covers detection through recovery only.

STEP 1 — Acknowledge alert, take IC
  Action: Acknowledge the PagerDuty alert (reward-worker queue depth > 10k) and assume IC role.
  Who: Bryce Harmon
  Verify: No explicit verification in thread (acknowledgment itself).
  Rollback: N/A — no state change.
  Trace: [M01]

STEP 2 — Measure queue depth
  Command: bundle exec rake sidekiq:queue_depth
  Who: Farid Osman
  Result in thread: reward queue at 48,213 pending jobs (normal: under 500).
  Verify: Command output confirms depth; 48,213 > 10k alert threshold matches the page.
  Rollback: N/A — read-only.
  Trace: [M02]

STEP 3 — Inspect dead set
  Action: Check the Sidekiq dead set contents.
  Who: Farid Osman
  Result in thread: 112 jobs, all Redis::TimeoutError, timestamped around 13:58.
  Command used: NOT STATED IN THREAD — needs confirmation before re-use.
  Verify: Dead set contents reported in thread; no command recorded.
  Rollback: N/A — read-only inspection.
  Trace: [M03]

STEP 4 — Pause enqueue (stop the bleed)
  Command: bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'
  Who: Farid Osman
  Verify: No explicit verification of the disable shown in thread — needs confirmation (e.g., no check recorded that the flag read as disabled).
  Rollback: bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)' (stated in-thread [M04]).
  State change: Yes — disables auto_recognition_enqueue.
  Trace: [M04]

STEP 5 — Clear the dead set
  Action: Cleared out the dead set from the console.
  Who: Elena Sinclair
  Command used: NOT STATED IN THREAD ("while I was in the console") — needs confirmation before re-use.
  Verify: No explicit post-clear verification in thread — needs confirmation.
  Rollback: NOT STATED IN THREAD — the 112 dead jobs were removed with no recorded backup/requeue path. Needs confirmation (e.g., whether they were retried or discarded).
  State change: Yes — dead set emptied.
  Trace: [M05]

STEP 6 — Scale workers up
  Command: kubectl scale deployment/reward-worker --replicas=6
  Who: Bryce Harmon
  Prior state: 3 replicas (stated in-thread).
  Verify: No explicit immediate verification in [M06]; downstream effect evidenced by drain in [M07] (depth 9,400, falling ~1,200/min within ~12 min).
  Rollback: kubectl scale deployment/reward-worker --replicas=3 (stated in-thread [M06]).
  State change: Yes — replica count 3 → 6.
  Trace: [M06]

STEP 7 — Monitor drain
  Action: Re-check queue depth and drain rate.
  Command: bundle exec rake sidekiq:queue_depth (per [M02]/[M08]; the exact command for the [M07] reading is not stated — needs confirmation).
  Who: Farid Osman
  Result in thread: depth 9,400, falling ~1,200/min.
  Verify: Depth decreasing toward 0.
  Rollback: N/A — read-only.
  Trace: [M07]

STEP 8 — Verify queue drained
  Command: bundle exec rake sidekiq:queue_depth
  Who: Cole Ingram
  Result in thread: returns 0; Datadog error rate back to baseline.
  Verify: Explicit — depth 0 and error rate at baseline.
  Rollback: N/A — read-only.
  Trace: [M08]

STEP 9 — Re-enable enqueue
  Command: bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'
  Who: Bryce Harmon
  Verify: Explicit — 40 new jobs processed cleanly in the following 3 minutes.
  Rollback: bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)' (inverse per [M04]).
  State change: Yes — re-enables auto_recognition_enqueue.
  Trace: [M09]

STEP 10 — Scale workers back down
  Command: kubectl scale deployment/reward-worker --replicas=3
  Who: Bryce Harmon
  Verify: Explicit in-thread — queue stable at 0; incident declared resolved.
  Rollback: kubectl scale deployment/reward-worker --replicas=6 (inverse per [M06]).
  State change: Yes — replica count 6 → 3.
  Trace: [M10]

GAPS / NEEDS CONFIRMATION (thread does not contain):
- Exact command for dead-set inspection [M03] and for clearing it [M05]; no rollback for the clear.
- Explicit verification that the enqueue-disable flag took effect at [M04].
- Root cause: only causal hint is Redis::TimeoutError on dead jobs from ~13:58 [M03]; no root-cause identification or fix step exists in the thread.
- The command used for the [M07] depth reading is not stated.
datadog-incident-timelinePASS 1.00
Plain-English question: Sequence the failure from the logs — what broke first, and what cascaded after it?
Prompt — what the model saw
From the log slice, sequence the failure: first error with timestamp, the cascade in order, the service and job involved, and the Datadog query that would confirm the first error. State what the logs do not show.
cost $0.0079 · 60s · in 618 / out 2,744 tok
✓ first_error_ts✓ service_named✓ cascade_ordered✓ query✓ gaps_stated

✓ Correct Answer (ground truth)

Plain-English answer: A reward-service Redis timeout at 14:01:12Z cascades to sidekiq at 14:01:40 and api-gateway at 14:03:05, with the Datadog query that confirms the first error.
The answer in one line: job_class_involved = RewardGiveJob
cascade_orderservice=reward-service; first_error_at=2026-09-03T14:01:12Z | service=sidekiq; first_error_at=2026-09-03T14:01:40Z | service=api-gateway; first_error_at=2026-09-03T14:03:05Z | service=web-app; first_error_at=2026-09-03T14:03:30Z
clean_servicespostgres
confirming_queryservice:reward-service status:error "@message:Redis::TimeoutError"
first_errortimestamp=2026-09-03T14:01:12Z; service=reward-service; message=Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s
job_class_involvedRewardGiveJob
logs_do_not_showroot cause of the Redis timeout, redis-server's own logs, infrastructure metrics (CPU/memory/network), any deploy or config change preceding the incident, who applied mitigation or when (no actor fields in logs)
graded checksfirst_error_ts · service_named · cascade_ordered · query · gaps_stated

Model's actual answer

show the model's full answer
**First error**
2026-09-03T14:01:12Z — reward-service — ERROR — "Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s"

**Cascade, in order** (all timestamps 2026-09-03Z)

1. 14:01:12 — reward-service: Redis connection timeout (origin).
2. 14:01:20 / 14:01:30 / 14:01:40 — reward-service: "retry exhausted for RewardGiveJob" (x3).
3. 14:01:40 — sidekiq: RewardGiveJob failed; retrying in 60s. Further sidekiq failures at 14:02:28, 14:03:31, 14:04:22, 14:05:26, 14:06:47.
4. 14:02:30 — sidekiq WARN: queue `reward` depth above 10,000.
5. 14:02:36–14:05:50 (sidekiq_jobs.csv) — collateral job class RecognitionDigestJob fails x4 (J-00013–J-00016), same Redis::TimeoutError.
6. 14:03:05 — api-gateway: 502 upstream timeout calling reward-service /gives; further 502s at 14:03:48, 14:04:13, 14:05:16, 14:06:52.
7. 14:03:30 — web-app: "Give form submission failed: upstream 502 from api-gateway"; further user-facing failures at 14:04:45, 14:05:42, 14:06:49.
8. Recovery: 14:22:10 — reward-service: "Redis connection restored; resuming job processing." 14:24:45 — sidekiq: queue depth below 500.

**Services and jobs involved**
- Services: reward-service (origin), sidekiq (worker/queue), api-gateway (edge), web-app (user-facing). Failing dependency is redis-primary — it appears only inside reward-service's error message, with no logs of its own.
- Jobs: RewardGiveJob — 12 failed (J-00001–J-00012); RecognitionDigestJob — 4 failed (J-00013–J-00016). Total 16, all Redis::TimeoutError.

**Arithmetic**
- Error burst in the log slice: 14:01:12 → 14:06:52 = 5 min 40 s of continuous errors.
- First error to Redis restored: 14:01:12 → 14:22:10 = 20 min 58 s.
- First error to queue drained below 500: 14:01:12 → 14:24:45 = 23 min 33 s.
- Job failures: 12 + 4 = 16.

**Datadog query confirming the first error**
`service:reward-service status:error "Redis::TimeoutError"` with time range 2026-09-03T14:00:00Z–14:02:00Z; the earliest hit should be 14:01:12Z with message "Connection to Redis at redis-primary:6379 timed out after 5s."

**What the logs do not show**
- Root cause of the Redis timeout — no redis-primary logs; network partition, memory pressure, failover, or deployment are all unconfirmed.
- Permanent user impact — whether any RewardGiveJob retries ultimately succeeded or died after final retry is not shown.
- No job IDs in the log slice (only in sidekiq_jobs.csv), and no linkage between J-IDs and specific log lines.
- A log gap from 14:06:52 to 14:10:56 with no reward-service/sidekiq/api-gateway/web-app entries — whether errors continued is unknown.
- No alerting, paging, or incident-declaration events.
- No revenue/dollar impact, customer aliases, or counts of affected end users.
- No deployment, config change, or rollback events preceding the failure.
feature-flag-state-summaryPASS 1.00
Plain-English question: Summarize what each feature flag controls and who has it on.
Prompt — what the model saw
Summarize the state of the listed feature flags: what each controls per the code excerpt, which companies or segments have it on, targeting rules, and any flag with no code reference or no targets. Company counts must match the export. Do not describe a flag that is not in the export.
cost $0.0065 · 39s · in 0 / out 2,120 tok
✓ orphans✓ no_targets✓ counts_match✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: Recognition streaks on for 42 beta companies, redeem-flow redesign off with a 12-company target list, plus any flag with no code reference or no targets.
The answer in one line: rule = company counts must equal the export; do not describe flags absent from the export.
code_referencedrecognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, redeem_flow_redesign, analytics_dashboard_v3, ms_teams_app_v2
flags{"recognition_streaks_v2": {"state": "on", "targeting_rules": "segment:beta_companies", "company_count": 42}, "points_budget_guardrails": {"state": "on", "targeting_rules": "all_companies", "company_c
flags_with_no_targetspaused_offboard_cleanup
orphan_flags_no_code_referencelegacy_give_modal, survey_boosters_q3, paused_offboard_cleanup
rulecompany counts must equal the export; do not describe flags absent from the export.
graded checksorphans · no_targets · counts_match · no_invented

Model's actual answer

show the model's full answer
Feature flag summary (9 flags in export; all described below)

1. recognition_streaks_v2
   - State: on
   - Targeting: segment:beta_companies
   - Companies: 42
   - Code: app/models/recognition.rb — when enabled, StreakTracker.record(give) runs on a recognition.

2. points_budget_guardrails
   - State: on
   - Targeting: all_companies
   - Companies: 220
   - Code: app/services/budget_service.rb — when enabled, BudgetService.new(company).enforce!(giver, points) runs.

3. slack_dm_nudges
   - State: on
   - Targeting: segment:region_na
   - Companies: 87
   - Code: app/jobs/nudge_job.rb — job exits early unless enabled; otherwise SlackDm.send_nudge(user) fires.

4. redeem_flow_redesign
   - State: off
   - Targeting: targeted_list
   - Companies: 12
   - Code: app/controllers/redeem_controller.rb — when enabled, renders RedeemV2Component; otherwise RedeemV1Component.
   - Note: state is off, but the export shows 12 companies under targeted_list.

5. analytics_dashboard_v3
   - State: on
   - Targeting: segment:tier_three
   - Companies: 65
   - Code: app/controllers/analytics_controller.rb — when enabled, @dashboard = AnalyticsV3.new(company).

6. ms_teams_app_v2
   - State: off
   - Targeting: targeted_list
   - Companies: 9
   - Code: app/services/teams_installer.rb — when enabled, TeamsAppV2.install(company) runs.
   - Note: state is off, but the export shows 9 companies under targeted_list.

7. legacy_give_modal
   - State: off
   - Targeting: segment:legacy_plan
   - Companies: 14
   - Code: NO code reference in the excerpt.

8. survey_boosters_q3
   - State: on
   - Targeting: segment:legacy_plan
   - Companies: 7
   - Code: NO code reference in the excerpt.

9. paused_offboard_cleanup
   - State: off
   - Targeting: none (blank)
   - Companies: 0
   - Code: NO code reference in the excerpt.
   - This is the only flag with no targets.

Flags with no code reference: legacy_give_modal, survey_boosters_q3, paused_offboard_cleanup (3 of 9).
Flags with no targets: paused_offboard_cleanup only.

Company count check (matches export exactly): 42 + 220 + 87 + 12 + 65 + 9 + 14 + 7 + 0 = 456 total flag-company rows.
nda-triagePASS 1.00
Plain-English question: Triage each NDA green, yellow, or red — and quote the clause that decides it.
Prompt — what the model saw
Triage each NDA as GREEN (standard approval), YELLOW (counsel review), or RED (full legal review). Cite the clause that drives each classification by quoting it. Catch embedded non-solicits, non-competes, missing carve-outs, and unusual term or governing law. One line of reasoning per NDA.
cost $0.0048 · 23s · in 0 / out 903 tok
✓ green✓ yellow✓ red✓ non_solicit_caught✓ clause_quoted✓ correct_labels

✓ Correct Answer (ground truth)

Plain-English answer: NDA-1 green (standard mutual terms), NDA-2 yellow for an embedded 18-month non-solicit, with the driving clause quoted for every classification.
The answer in one line: non_solicit_in = NDA-2
checksclassification_per_nda, clause_quotes_exist_in_document, non_solicit_caught, carveouts_checked
classifications{"NDA-1": {"classification": "GREEN", "driving_reason": "standard mutual NDA: 2-year term, standard carve-outs, Delaware law", "driving_clause_quote": "Carve-outs: information that (a) is or becomes p
non_solicit_inNDA-2
graded checksgreen · yellow · red · non_solicit_caught · clause_quoted · correct_labels

Model's actual answer

show the model's full answer
NDA-1 — GREEN: Standard mutual NDA; full carve-out set ("(a) is or becomes publicly available through no breach... (e) must be disclosed by law or court order"), benign term (2 yrs + 3-yr survival = max 5 yrs total, arithmetic: 2+3), Delaware law, and "No license, no obligation to proceed, no exclusivity" — nothing to escalate.

NDA-2 — YELLOW: Standard terms otherwise (3-yr term, standard carve-outs, Delaware law), but the embedded non-solicit — "During the term of this Agreement and for eighteen (18) months thereafter, neither party shall solicit for employment or hire any employee of the other party with whom it came into contact under this Agreement" — restricts hiring for 3 yrs + 18 mo = 4.5 yrs total (arithmetic: 3+1.5); it's mutual and contact-limited so counsel review (not full legal) suffices.

NDA-3 — RED: Multiple triggers — (1) the non-compete "For a period of three (3) years following the Effective Date, Recipient shall not, directly or indirectly, engage in or provide services to any business that competes with Discloser's business" bars Bonusly's core business with no field/geographic carve-out; (2) "No confidentiality carve-outs are stated in this Agreement" means even public info stays restricted; (3) one-way confidentiality in favor of [PARTY C] plus an unusual 5-yr term; and (4) foreign governing law — "the laws of the Republic of Ireland, and the parties submit to the exclusive jurisdiction of its courts" — requiring full legal review.