← back to all models ← RevenueBench home
xiaomi

xiaomi/mimo-v2.6-pro

Mean 0.984 · 35/38 perfect tests · $0.62 total · 38/40 tests completed — remaining tests discarded from study (endpoint timeouts after dedicated retries) · modalities in: text, image, video, audio · out: text · each card: the prompt → the correct answer (gold) → this model's actual answer

What do the modalities mean?

Modalities in means what you can send this model: text, images, files, video, or audio. Modalities out means what it can send back. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what this model can handle in real work.

data-crm 4 tests
owner-resolutionPASS 1.00
Plain-English question: Every open deal has a blank owner name — resolve each one to a real person using the owners table.
Prompt — what the model saw
The deal extract below has an owner name field that is blank on nearly every row and an owner id field that is populated. Using the owners table, resolve each open deal to an owner name. Report: number of deals resolved, deals whose owner id has no match in the owners table (list them), the owner ids that map to archived/deactivated owners (list their names), and the total pipeline amount per resolved owner. Do not guess a name for any unmatched id.
cost $0.0360 · 501s · in 64,661 / out 8,946 tok
✓ resolved_count✓ unmatched_none✓ deactivated_flagged✓ jake_amount✓ justin_amount✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: All 156 deals resolve to six named owners; jake_amount maps to Bryce Harmon ($1,054,144) and justin_amount to Alex Franklin ($624,310).
The answer in one line: jake_amount = 1,054,144; justin_amount = 624,310; resolved = 156
jake_amount1,054,144
justin_amount624,310
resolved156
deactivated
per_ownerBryce Harmon=1,054,144; Alex Franklin=624,310; Dana Mercer=341,195; Cole Ingram=288,161.43; Farid Osman=4,134; Elena Sinclair=2,100
graded checksresolved_count · unmatched_none · deactivated_flagged · jake_amount · justin_amount · no_invented

Model's actual answer

show the model's full answer
Deals resolved to owner names: 156 of 156 (every owner_id in deals_open.csv matches a row in owners.csv).

Deals whose owner_id has no match in the owners table: none. No deal is left unresolved, so there is nothing to list.

Owner ids in the extract that map to archived/deactivated owners: none. The owners table contains two archived owners — 1520255671 (Gavin Porter, archived=true) and 77260721 (Hugo Lindqvist, archived=true) — but no open deal carries either id, so no open deal sits with an archived owner. (They are listed here for completeness only; they are not deal owners in this extract.)

Total pipeline amount per resolved owner (sum of amount over that owner's open deals):

Bryce Harmon (119337721) — 35 deals
  24,000 + 19,656 + 13,500 + 7,000 + 2,520 + 240,000 + 99,000 + 72,000 + 70,000
  + 63,600 + 45,000 + 1 + 21,000 + 23,400 + 13,680 + 5,502 + 8,160 + 1 + 11,400
  + 1 + 36,000 + 31,500 + 6,000 + 10,800 + 30,275 + 17,400 + 12,600 + 18,000
  + 37,440 + 18,828 + 2,880 + 36,000 + 20,880 + 10,920 + 25,200 = 1,054,144.00

Alex Franklin (84342457) — 67 deals
  14,850 + 13,770 + 11,200 + 9,000 + 6,360 + 5,400 + 3,240 + 2,484 + 1,920
  + 1,080 + 7,200 + 19,000 + 2,880 + 1,400 + 4,800 + 1,632 + 10,000 + 9,300
  + 2,700 + 2,160 + 1,800 + 3,600 + 3,840 + 15,000 + 1,968 + 4,000 + 3,600
  + 4,800 + 3,120 + 2,520 + 9,000 + 2,400 + 62,000 + 5,400 + 5,100 + 16,700
  + 4,400 + 1,620 + 2,600 + 7,200 + 18,000 + 17,000 + 8,316 + 8,100 + 18,000
  + 12,600 + 24,000 + 15,000 + 9,000 + 7,200 + 3,780 + 16,200 + 7,200 + 4,680
  + 1,800 + 18,000 + 2,730 + 2,400 + 3,060 + 18,000 + 12,000 + 1,800 + 4,400
  + 31,200 + 7,200 + 1,600 + 60,000 = 624,310.00

Dana Mercer (83155923) — 24 deals
  11,250 + 10,500 + 9,000 + 9,000 + 5,400 + 4,800 + 4,600 + 1,920 + 15,000
  + 4,200 + 18,900 + 27,000 + 43,875 + 20,000 + 60,000 + 8,100 + 16,250
  + 3,150 + 5,000 + 2,100 + 23,400 + 5,400 + 7,350 + 25,000 = 341,195.00

Cole Ingram (83155924) — 22 deals
  2,760 + 1,330 + 7,225.40 + 3,360 + 5,616 + 2,700 + 7,781.20 + 6,947.50
  + 40,000 + 12,168 + 2,480.40 + 11,116 + 11,193 + 1,875 + 31,750 + 32,175
  + 4,779.88 + 3,334.80 + 58,529.25 + 4,140 + 18,000 + 18,900 = 288,161.43

Farid Osman (716654662) — 7 deals
  1,249 + 1 + 1 + 1 + 2,880 + 1 + 1 = 4,134.00

Elena Sinclair (701163055) — 1 deal
  2,100 = 2,100.00

Cross-check: 35 + 67 + 24 + 22 + 7 + 1 = 156 deals.
1,054,144.00 + 624,310.00 + 341,195.00 + 288,161.43 + 4,134.00 + 2,100.00 = 2,314,044.43 total open pipeline.
gong-hubspot-join-integrityPASS 1.00
Plain-English question: How many call records point at deal ids that do not exist in the open pipeline?
Prompt — what the model saw
Check the integrity of the call-to-deal mapping below. The gong table lists calls in the last 90 days keyed to HubSpot deal ids; deals_open lists the current open default-pipeline deals. Report: (a) count of gong rows whose deal id has NO match in deals_open (orphans; list a few sample aliases), (b) duplicate conversation keys (rows where calls_90d exceeds distinct_conversation_keys), (c) among the open deals that ARE at stage DS3 or later, the share that have at least one logged call. Give sample aliases for the orphan finding, all taken from the data. Note: gong_owner_id is a Gong user id, not a HubSpot owner id.
cost $0.0477 · 553s · in 93,929 / out 7,852 tok
✓ orphan_count✓ dup_keys_0✓ ds3_coverage✓ sample_ids_present

✓ Correct Answer (ground truth)

Plain-English answer: 30 Gong call rows are orphans keyed to dead deal ids, none are duplicates, and DS3 call coverage is 29.4%.
The answer in one line: orphans = 30; ds3_cov_pct = 29.4; dup = 0
orphans30
ds3_cov_pct29.4
dup0
orphan_aliasesDeal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42, Deal-36EA09
graded checksorphan_count · dup_keys_0 · ds3_coverage · sample_ids_present

Model's actual answer

show the model's full answer
CALL-TO-DEAL MAPPING INTEGRITY CHECK

(a) ORPHANS — gong rows with deal id NOT in deals_open
Count: 30 of 67 gong rows (30/67 = 44.8%).
Arithmetic: 67 gong rows total − 37 rows whose hs_deal_id appears in deals_open (156 rows) = 30 orphans.
Sample orphan aliases (all from gong_calls_by_deal_90d.csv):
  Deal-8FA85D (60251290957, 46 calls)
  Deal-8FC3F9 (60251649055, 24 calls)
  Deal-3B7945 (60251639682, 21 calls)
  Deal-42B265 (61227242540, 21 calls)
  Deal-9CCC42 (61430316324, 17 calls)
  Deal-36EA09 (60251082126, 17 calls)
  Deal-9A43B4 (60250446726, 14 calls)
  Deal-605F3C (60251622924, 12 calls)
  Deal-E2D34B (60251652002, 11 calls)
  Deal-76821A (60251092117, 11 calls)
Note: several orphans carry last_call dates into Oct–Nov 2026 (e.g. Deal-42B265 last_call 2026-11-25, Deal-9CCC42 2026-11-27) despite a "90 days" label — the window definition and the deal-id mapping both warrant a look.

(b) DUPLICATE CONVERSATION KEYS — rows where calls_90d > distinct_conversation_keys
Count: 0. Every one of the 67 rows has calls_90d == distinct_conversation_keys (checked row by row, e.g. 60251290957: 46 = 46; 60333965090: 11 = 11; 61749892214: 4 = 4). No duplicated conversation keys detected in this extract.

(c) DS3+ OPEN DEALS WITH ≥1 LOGGED CALL
Denominator: open deals at stage DS3, DS4, or DS5 = 85 deals (of 156 open).
  DS3 = 49, DS4 = 22, DS5 = 14 → 49 + 22 + 14 = 85.
Numerator: of those 85, 25 have a hs_deal_id present in the gong table (≥1 call in the window).
Share: 25 / 85 = 29.4%.
The remaining 60 (70.6%) DS3+ deals have no gong call rows at all — e.g. Deal-9AAE5F (DS4), Deal-403845 (DS5), Deal-2465CE (DS5), Deal-FC22A3 (DS3), Deal-7BBDFA (DS3).

Summary: mapping integrity is weak — 44.8% of gong call rows point at deal ids absent from the open-deal extract (orphans concentrated in the 6025… / 61xx… id range), though conversation-key duplication is clean; and call coverage of late-stage pipeline is low at 29.4%.
snowflake-sql-trapped-catalogPASS 1.00
Plain-English question: Write one Snowflake query for first-month usage per customer using only the provided, trap-laden catalog.
Prompt — what the model saw
Using ONLY the catalog provided, write one Snowflake SQL query returning, per customer company, for its first calendar month as a customer: unique givers, recognition count, successful redemption count. Read the catalog notes carefully: some tables are stale or unpopulated and must be avoided, and one documented filter (excluding deleted givers) must NOT be applied to historical giving counts. Put a comment above each FROM or JOIN explaining the table choice. Return only fenced SQL.
cost $0.0117 · 414s · in 20,302 / out 3,305 tok
✓ avoids_stale_deals✓ avoids_unpopulated_live✓ uses_cohort_summary✓ succeeded_filter✓ no_giver_deletion_filter✓ returns_sql

✓ Correct Answer (ground truth)

Plain-English answer: A single correct SQL query that uses the right catalog tables and avoids the trap columns — graded on the exact tables and columns.
The answer in one line: use_table = PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
avoid_stalePRODUCTION.HUBSPOT.DEALS (stale)
avoid_unpopulatedV2_LIVE.OBJECTS_DEALS (unpopulated)
forbidden_filterdeleted-giver exclusion must NOT be applied
outputone fenced SELECT with per-FROM/JOIN table-choice comment
required_filtersucceeded redemptions only
use_tablePRODUCTION.PLG.COMPANY_COHORT_SUMMARY
graded checksavoids_stale_deals · avoids_unpopulated_live · uses_cohort_summary · succeeded_filter · no_giver_deletion_filter · returns_sql

Model's actual answer

show the model's full answer
```sql
/*
  Per customer company -- first calendar month as a customer:
  unique givers, recognition count, successful redemption count.

  EXPLICIT DATA GAPS IN THE CATALOG (stated, not invented):
  1. No raw giving/recognition event table exists in this catalog. The only
     per-company giving figures available are the pre-aggregated M1_* columns
     (M1_USERS, M1_GIVES) on COMPANY_COHORT_SUMMARY, defined as first month as
     a customer. If SIGNUP_DATE is mid-month, M1_* (month-1-from-signup) is not
     exactly the first CALENDAR month, and nothing in this catalog lets us
     recompute a calendar-month giver/recognition count.
  2. REDEMPTION_RECORDS_V2 is the documented redemption source, but the catalog
     documents ONLY its STATE column -- no company key or event-date columns are
     listed. The join below uses ASSUMED column names (COMPANY_ID, REDEEMED_AT);
     these MUST be confirmed against the real schema before running.
  3. COMPANY_COHORT_SUMMARY's company identifier column is also not named in the
     catalog; COMPANY_ID is assumed below and must be confirmed.
  4. Tables deliberately AVOIDED: HUBSPOT_HUB_1973303.V2_LIVE.OBJECTS_DEALS
     (unpopulated), PRODUCTION.HUBSPOT.DEALS (stale, last sync 2023-03),
     PRODUCTION.HUBSPOT.GONG_HUBSPOT_MAP_FAST (retired pattern).
  5. The documented "exclude deleted givers" filter is deliberately NOT applied
     anywhere here -- per the catalog it understates historical giving counts.
*/

SELECT
    c.COMPANY_ID,
    DATE_TRUNC('month', c.SIGNUP_DATE)::DATE        AS first_customer_month,
    c.M1_USERS                                      AS unique_givers,
    c.M1_GIVES                                      AS recognition_count,
    COUNT_IF(r.STATE = 'succeeded')                 AS successful_redemptions
-- FROM: PRODUCTION.PLG.COMPANY_COHORT_SUMMARY is the only table in the catalog
-- with per-customer-company, first-month-as-a-customer giving metrics (one row
-- per self-serve company; SIGNUP_DATE anchors the first calendar month and
-- M1_USERS/M1_GIVES supply unique givers and recognition count, since no raw
-- giving table exists). No deleted-giver filter applied, per business rules.
FROM PRODUCTION.PLG.COMPANY_COHORT_SUMMARY AS c
-- JOIN: PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 is the
-- documented source for redemption counts (despite the DEPRECATED schema name),
-- and the only place the STATE = 'succeeded' rule can be applied. Restricted to
-- events inside the company's first calendar month. Columns beyond STATE are
-- undocumented -- COMPANY_ID and REDEEMED_AT here are ASSUMPTIONS to verify.
LEFT JOIN PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 AS r
       ON r.COMPANY_ID = c.COMPANY_ID
      AND r.REDEEMED_AT >= DATE_TRUNC('month', c.SIGNUP_DATE)
      AND r.REDEEMED_AT <  DATE_TRUNC('month', c.SIGNUP_DATE) + INTERVAL '1 month'
GROUP BY
    c.COMPANY_ID,
    DATE_TRUNC('month', c.SIGNUP_DATE)::DATE,
    c.M1_USERS,
    c.M1_GIVES
ORDER BY
    first_customer_month,
    c.COMPANY_ID;
```
crm-completeness-enrichmentPASS 1.00
Plain-English question: Audit the CRM for missing and invalid fields, and propose the fixes.
Prompt — what the model saw
Audit the CRM extract for completeness and propose fixes. Deals need owner, stage, amount, close date, why-buys; companies need industry, employee count, HQ country; contacts need email, title, persona. Report completeness percent per field. Identify duplicate company clusters (name variants, shared domain) with a survivor per cluster. List invalid emails and domain mismatches. Fill missing company fields only where the enrichment export has a matching row; where CRM and enrichment disagree list both and recommend a source. Never invent a value. End with the 10 fixes with the most pipeline amount at stake.
cost $0.0131 · 174s · in 6,164 / out 11,913 tok
✓ dup_clusters✓ invalid_emails✓ survivor_language✓ completeness_pct✓ no_invent

✓ Correct Answer (ground truth)

Plain-English answer: 9 companies missing employee count, 6 missing HQ country, 4 invalid emails, 1 domain mismatch, 2 duplicate clusters, and 25 enrichment matches.
The answer in one line: n_contacts = 52; n_companies = 34; enrichment_matches = 25; invalid_email_count = 4; n_dup_clusters = 2; domain_mismatch_count = 1 (+4 more below)
n_contacts52
domain_mismatch_count1
enrichment_matches25
invalid_email_count4
n_companies34
n_dup_clusters2
company_missingindustry=0; employee_count=9; hq_country=6
domain_mismatchesCT-0011
duplicate_clusters['C-0A092931', 'C-0A092932'], ['C-0A092933', 'C-0A092934']
invalid_emailsCT-0010, CT-0080, CT-0081, CT-0192
graded checksdup_clusters · invalid_emails · survivor_language · completeness_pct · no_invent

Model's actual answer

show the model's full answer
DATA AUDIT — CRM EXTRACT
Sources used: companies.csv (34 rows), contacts.csv (52 rows), zoominfo_enrichment.csv (25 rows).
MISSING DATA: no deals.csv was provided. Deals fields (owner, stage, amount, close date, why-buys) cannot be audited and NO pipeline amount can be computed. Consequently the final "10 fixes by pipeline at stake" CANNOT be pipeline-ranked — see section 7, ranked by records affected instead, clearly labeled. No amounts were invented.

────────────────────────────────────────
1. COMPLETENESS BY FIELD
────────────────────────────────────────
COMPANIES (n = 34)
  industry ............ 34/34 = 100.0%   (all populated; 10 dirty values — see 5)
  employee_count ...... 24/34 =  70.6%   (10 blank: C-EC3025, C-96039F, C-44EA29, C-D04904,
                                          C-B23205, C-60C75F, C-2C60E5, C-7BBDFA, C-50D386, C-93C8BF)
  hq_country .......... 28/34 =  82.4%   (6 blank: C-2D1F1B, C-D73B89, C-44EA29, C-D04904,
                                          C-2C60E5, C-EE9FFB)

CONTACTS (n = 52)
  email ............... 48/52 =  92.3% valid (52/52 populated, but 4 unparseable — see 3)
  title ............... 39/52 =  75.0% (13 blank: CT-0000, CT-0022, CT-0072, CT-0080, CT-0081,
                                         CT-0092, CT-0120, CT-0121, CT-0122, CT-0132,
                                         CT-0141, CT-0162, CT-0170)
  persona ............. 37/52 =  71.2% (15 blank: CT-0000, CT-0022, CT-0041, CT-0060, CT-0070,
                                         CT-0081, CT-0082, CT-0092, CT-0110, CT-0132,
                                         CT-0162, CT-0171, CT-0172, CT-0180, CT-0181)

DEALS — n = 0 (file absent). owner / stage / amount / close date / why-buys: NOT AUDITABLE.

Note: there is no contact enrichment file, so blank titles/personas have NO fill source in the data provided — they must be collected, not filled.

────────────────────────────────────────
2. DUPLICATE COMPANY CLUSTERS
────────────────────────────────────────
Cluster A — shared domain acme-corp.com (no name column exists; aliases are opaque hashes, so only domain-based dupes are detectable)
  C-0A092931  acme-corp.com  Technology  500  US
  C-0A092932  acme-corp.com  tech        510  USA
  SURVIVOR: C-0A092931 (lowest alias ID = first-created; canonical industry spelling "Technology").
  Merge-safe fields: industry (tech → Technology, same category), country (US = USA = United States).
  UNRESOLVED CONFLICT: employee_count 500 vs 510. zoominfo_enrichment.csv has NO acme-corp.com row → cannot arbitrate. Do not pick a value; flag for manual verification.

Cluster B — shared domain globex.io
  C-0A092933  globex.io  SaaS         200  US
  C-0A092934  globex.io  Technology   200  US
  SURVIVOR: C-0A092933 (lowest alias ID). employee_count agrees (200 = 200), country agrees.
  UNRESOLVED CONFLICT: industry SaaS vs Technology. No globex.io row in enrichment → flag, do not guess.

No other clusters: the 30 hash-alias companies each have a unique domain; no name field exists to detect name-variant dupes.

────────────────────────────────────────
3. INVALID EMAILS (4)
────────────────────────────────────────
  CT-0010  user0@          — empty domain after "@" (company C-66D1FC, 66d1fc.com)
  CT-0080  user0@          — empty domain after "@" (company C-92D97D, 92d97d.com)
  CT-0081  user1@          — empty domain after "@" (company C-92D97D, 92d97d.com)
  CT-0192  user2@          — empty domain after "@" (company C-425E2A, 425e2a.com)

DOMAIN MISMATCH (1)
  CT-0011  user1@other-domain.com vs contact domain 66d1fc.com vs company C-66D1FC domain 66d1fc.com
  → email domain (other-domain.com) ≠ company domain. Do not auto-correct the local part; verify ownership, then either re-point email to @66d1fc.com or reassign the contact's company.

All other 47 valid emails match their company domain.

────────────────────────────────────────
4. ENRICHMENT FILLS (missing CRM value, ZI row exists) — 9 fills, all employee_count
────────────────────────────────────────
  C-EC3025   ec3025.com   employee_count '' → 400
  C-96039F   96039f.com   employee_count '' → 400
  C-44EA29   44ea29.com   employee_count '' → 400
  C-D04904   d04904.com   employee_count '' → 400
  C-B23205   b23205.com   employee_count '' → 400
  C-60C75F   60c75f.com   employee_count '' → 400
  C-2C60E5   2c60e5.com   employee_count '' → 340
  C-7BBDFA   7bbdfa.com   employee_count '' → 400
  C-50D386   50d386.com   employee_count '' → 400

  Post-fill: employee_count 24 + 9 = 33/34 = 97.1%. Only C-93C8BF (93c8bf.com) stays blank — no enrichment row exists.

  NOT fillable (ZI row exists but ZI value is also blank — not a disagreement, just absent in both):
  hq_country for C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5.
  NOT fillable (no ZI row at all): C-BA969B, C-332637, C-93C8BF, C-EE9FFB, C-C9BB20,
  C-0A092931, C-0A092932, C-0A092933, C-0A092934 (9 of 34 = 26.5% uncovered by enrichment).

────────────────────────────────────────
5. CRM vs ENRICHMENT DISAGREEMENTS (both populated, different)
────────────────────────────────────────
A) INDUSTRY — 10 rows. Same category, different vocabulary (CRM "Technology/tech/Tech " vs ZI "Computer Software"):
  C-66D1FC  tech        | Computer Software
  C-EC3025  Technology  | Computer Software
  C-44EA29  tech        | Computer Software
  C-92D97D  Technology  | Computer Software
  C-D04904  Technology  | Computer Software
  C-77A95A  Technology  | Computer Software
  C-AA8DDA  Technology  | Computer Software
  C-B25F40  Technology  | Computer Software
  C-60C75F  tech        | Computer Software
  C-425E2A  "Tech "     | Computer Software
  RECOMMEND SOURCE: ZoomInfo. It is a single external taxonomy applied consistently across all 25 enrichment rows; CRM holds three spellings of the same value. This is a vocabulary mismatch, not a factual conflict — adopting one vocabulary is the fix.

B) EMPLOYEE_COUNT — 1 conflict, inside duplicate cluster A only:
  C-0A092931 500 vs C-0A092932 510 (no ZI row) → UNRESOLVED, manual verification required. No source recommended; picking one would be inventing a fact.

C) HQ_COUNTRY — 10 rows differ only in format (US / USA vs "United States"): C-66D1FC, C-950043, C-77A95A, C-D0662E, C-B23205, C-EC3025, C-96039F, C-E51FB7, C-425E2A, C-2D7423.
  RECOMMEND SOURCE: ZoomInfo ("United States") as the canonical string. All agree on the actual country — no factual conflicts.

D) OTHER NORMALIZATION (no ZI involvement):
  "Tech " trailing whitespace: C-425E2A, C-BA969B, C-93C8BF, C-C9BB20.
  "health care" vs "Healthcare": C-7BBDFA, C-50D386 vs C-B23205, C-2C60E5, C-63A874, C-EE9FFB. (ZI also says "health care" for 7bbdfa/50d386 — agree with each other, but the CRM file itself uses two spellings.)

────────────────────────────────────────
6. WHAT CANNOT BE FIXED FROM THIS DATA (explicitly)
────────────────────────────────────────
  - All deal fields (no deals.csv at all).
  - 13 blank titles + 15 blank personas (no contact enrichment source provided).
  - 4 invalid emails (local parts unknown — cannot reconstruct user0@/user1@/user2@ addresses).
  - 6 blank hq_country values (5 blank in both sources, 1 with no source row).
  - C-93C8BF employee_count (no enrichment row).
  - Cluster A/B field conflicts (500 vs 510; SaaS vs Technology) — no arbitrating source.

────────────────────────────────────────
7. TOP 10 FIXES
────────────────────────────────────────
WARNING: pipeline-at-stake ranking is IMPOSSIBLE — no deals.csv, no amounts exist in the data provided. Below is ranked by RECORDS AFFECTED instead. Re-rank once deals data lands.

  #  Fix                                          Records  Pipeline $ at stake
  1  Industry taxonomy: normalize tech/tech␣/      20       UNKNOWN (no deals data)
     health care + adopt ZI vocabulary
  2  HQ country normalize US/USA → "United States" 16       UNKNOWN
  3  Fill 9 employee_count values from ZoomInfo      9       UNKNOWN
  4  Collect 15 missing personas (no source)        15       UNKNOWN
  5  Collect 13 missing titles (no source)          13       UNKNOWN
  6  Fix 4 invalid emails (CT-0010/0080/0081/0192)   4       UNKNOWN
  7  Merge C-0A092932 → survivor C-0A092931         2→1     UNKNOWN
  8  Merge C-0A092934 → survivor C-0A092933         2→1     UNKNOWN
  9  Resolve CT-0011 email/domain mismatch           1      UNKNOWN
 10  Manual research: 6 blank hq_country +          7        UNKNOWN
     C-93C8BF employee_count (no enrichment rows)

Also queued but sub-threshold: 4 trailing-space industry values ("Tech ") fall under fix #1.

Bottom line: company fields are 70–100% complete and 97.1% recoverable on employee_count via enrichment; contacts are 71–92% complete with NO fill source; deals are entirely absent and block the pipeline-weighted prioritization you asked for.
deal-intelligence 4 tests
pipeline-tieringPASS 1.00
Plain-English question: Score all 156 open deals into exactly one of LOCK / ACTION / BUILD / REVIVE / WATCH / RISKY.
Prompt — what the model saw
Score the open pipeline below on these signals: forecast category, stage, meetings_30d, contact count, and engagement recency. Assign exactly one tier per deal: LOCK, ACTION, BUILD, REVIVE, WATCH, or RISKY (RISKY = the forecast category disagrees with the engagement evidence). Never tier a deal LOCK if it has zero meetings_30d. Do NOT list every deal. Return ONLY a compact JSON object with exactly these keys: - "tier_counts": object mapping each tier to the number of deals in it (must sum to the total deal count), - "tier_examples": object mapping each tier to up to 3 example deal aliases from the data, - "risky_deals": array of the deal aliases you tiered RISKY, - "lock_violations": number of deals tiered LOCK that have zero meetings_30d (should be 0), - "pipeline_shape": one short paragraph on the overall pipeline shape. Note: inbound_emails_30d is 0 for every row due to a data defect, so treat meetings_30d as the inbound signal.
cost $0.0782 · 589s · in 81,001 / out 46,026 tok
✓ json_returned✓ tier_vocabulary✓ tier_counts_sum✓ risky_deals_listed✓ no_lock_violation_claimed

✓ Correct Answer (ground truth)

Plain-English answer: All 156 open deals tiered exactly once across the six tiers using forecast category, stage, meetings, contacts, and recency.
The answer in one line: total = 156
total156
graded checksjson_returned · tier_vocabulary · tier_counts_sum · risky_deals_listed · no_lock_violation_claimed

Model's actual answer

show the model's full answer
{
  "tier_counts": {"LOCK": 3, "ACTION": 27, "BUILD": 60, "REVIVE": 1, "WATCH": 61, "RISKY": 4},
  "tier_examples": {"LOCK": ["Deal-D348E1", "Deal-C26D20", "Deal-403845"], "ACTION": ["Deal-25F752", "Deal-E53952", "Deal-547B2B"], "BUILD": ["Deal-499BF6", "Deal-C6FE92", "Deal-D73B89"], "REVIVE": ["Deal-2D1F1B"], "WATCH": ["Deal-66D1FC", "Deal-950043", "Deal-B23205"], "RISKY": ["Deal-A5E80A", "Deal-7BBDFA", "Deal-4A13AD"]},
  "risky_deals": ["Deal-A5E80A", "Deal-7BBDFA", "Deal-4A13AD", "Deal-690476"],
  "lock_violations": 0,
  "pipeline_shape": "Counts sum 3+27+60+1+61+4 = 156 deals ($2,314,044 total). Tiering arithmetic (data as-of 2026-09-04): RISKY = COMMIT/BEST_CASE where the category conflicts with engagement evidence (DS1 stage, or zero meetings_30d on DS1/DS2, or zero meetings_30d and >21 days since last touch); LOCK = COMMIT + DS4/DS5 + meetings_30d>=1 + touch <=21d + n_contacts>=3 (3 deals, $36,270, so no LOCK has zero meetings); REVIVE = >30 days since any touch; ACTION = fresh DS4/DS5 (touch <=21d) or meetings_30d>=1 with close date within -25/+30 days of as-of; BUILD = active DS3 (meetings or touch <=14d) or DS2 with meetings; WATCH = the quiet remainder. Shape: the pipeline is bottom-heavy and long-dated — 92 of 156 deals sit at DS1/DS2 (mostly PIPELINE, 0 meetings_30d, quiet), while just 3 DS5 COMMIT deals clear the LOCK bar; near-term money is concentrated in the 27 ACTION deals (21 of which are DS4/DS5 with recent touches but several have zero meetings_30d — e.g. Deal-547B2B, Deal-B7EBD1, Deal-A2B47C, Deal-2465CE are COMMIT yet meeting-less, kept at ACTION not LOCK per the rule); the 4 RISKY deals total $45,720 of over-claimed forecast (Deal-A5E80A is COMMIT at DS1; Deal-7BBDFA is BEST_CASE with 0 meetings and a 2026-07-21 last touch; Deal-4A13AD BEST_CASE, 0 meetings, 25d stale; Deal-690476 BEST_CASE at DS2 with 0 meetings). Data gaps: Deal-3EED2C and Deal-57FF13 have no row in engagements_by_deal_90d.csv (and Deal-57FF13's deals_open.csv row is truncated with missing last_contacted_field/n_contacts), so their engagement signals could not be verified — both tiered WATCH, not LOCK. Also noted as given: inbound_emails_30d is 0 on every engagement row (defect), so meetings_30d served as the inbound signal throughout."
}
call-transcript-extractionPASS 1.00
Plain-English question: Extract the CRM write-back fields from each call transcript as JSON.
Prompt — what the model saw
For each transcript, extract CRM write-back fields as JSON: why-buys (prospect statements only), pain points, stakeholders from the speaker list, budget signal (prospect-stated or null), timeline signal, competitor mentioned (only if the prospect raised it), next step (only if explicitly agreed), objections, confidence. Never fill a field from what the rep said or hoped.
cost $0.0047 · 48s · in 5,049 / out 2,805 tok
✓ json_emitted✓ budget_signal✓ null_discipline✓ stakeholders✓ competitor_prospect_only

✓ Correct Answer (ground truth)

Plain-English answer: Per transcript: prospect-only why-buys, pain points, stakeholders, budget and timeline signals, and competitors — e.g. TX-001: '$40k earmarked', live before open enrollment.
The answer in one line: rep = Alex Franklin
allowed_newTX-001, TX-002, TX-003, TX-004, TX-005, TX-006, Achievers, Bucketlist, Nectar
per_transcript{"TX-001": {"why_buys": ["automating anniversary and birthday awards"], "pain_points": ["HR team of three cannot keep up with it manually", "track everything in a spreadsheet, and people slip through
repAlex Franklin
rep_sourced_nulls{"TX-003": ["budget_signal \u2014 only the rep mentioned pricing/budget"], "TX-004": ["next_step \u2014 rep proposed, prospect said 'no promises'"], "TX-006": ["budget_signal \u2014 rep decoy, prospec
graded checksjson_emitted · budget_signal · null_discipline · stakeholders · competitor_prospect_only

Model's actual answer

show the model's full answer
Below: one JSON object per transcript. All values drawn only from the CSV; rep statements (Alex Franklin) excluded from why-buys, budget, timeline, competitor.

TX-001 — Deal-CFE7F4
```json
{
  "transcript_id": "TX-001",
  "deal_alias": "Deal-CFE7F4",
  "why_buys": ["Automating anniversary and birthday awards — HR team of three cannot keep up manually"],
  "pain_points": ["HR team of three cannot keep up with manual anniversary/birthday awards", "Tracking in a spreadsheet, people slip through the cracks"],
  "stakeholders": ["Alex Franklin (rep)", "Prospect (VP People)", "Prospect (HR Admin)"],
  "budget_signal": "About $40k earmarked for engagement tools this fiscal year",
  "timeline_signal": "Live before open enrollment in November",
  "competitor_mentioned": "Achievers (looked at last year, too heavy for a team their size)",
  "next_step": "Security review with IT lead on September 12",
  "objections": ["Need SSO and audit logs for IT to sign off"],
  "confidence": "high"
}
```

TX-002 — Deal-70BB30
```json
{
  "transcript_id": "TX-002",
  "deal_alias": "Deal-70BB30",
  "why_buys": ["Tie recognition to retention for their hourly workforce"],
  "pain_points": ["Regretted turnover over 30% in hourly workforce"],
  "stakeholders": ["Alex Franklin (rep)", "Prospect (Head of Total Rewards)", "Prospect (CFO)"],
  "budget_signal": "$25k pilot budget approved by finance for this quarter",
  "timeline_signal": "Decision by end of September",
  "competitor_mentioned": null,
  "next_step": "Send pilot agreement; prospect routes it to legal this week",
  "objections": ["Integration with Workday has to be rock solid — CFO's one condition"],
  "confidence": "high"
}
```

TX-003 — Deal-530B50
```json
{
  "transcript_id": "TX-003",
  "deal_alias": "Deal-530B50",
  "why_buys": ["Make recognition visible across their 12 retail locations"],
  "pain_points": ["Store managers have zero budget autonomy for on-the-spot recognition"],
  "stakeholders": ["Alex Franklin (rep)", "Prospect (People Ops Manager)"],
  "budget_signal": null,
  "timeline_signal": "No rush until Q1",
  "competitor_mentioned": "Bucketlist (CEO used it at her last company and liked it)",
  "next_step": "Schedule a call with the CEO; People Ops Manager to send two times",
  "objections": ["CEO must be sold first — she decides anything people-related"],
  "confidence": "medium"
}
```
Note: the CEO is referenced as decision maker but is not in the speaker list, so she is not listed as a stakeholder per the stated rule. Budget = null: the only pricing figure ("$8 per employee per month") was stated by the rep, and the prospect gave no figure or approval amount.

TX-004 — Deal-180D02
```json
{
  "transcript_id": "TX-004",
  "deal_alias": "Deal-180D02",
  "why_buys": ["Consolidate three separate recognition tools into one"],
  "pain_points": ["Paying for three tools", "None of the three tools talk to their HRIS"],
  "stakeholders": ["Alex Franklin (rep)", "Prospect (VP People)", "Prospect (IT Security Lead)"],
  "budget_signal": "Under $15k annually, VP People can approve without going to the board",
  "timeline_signal": "Procurement cycle runs six to eight weeks minimum",
  "competitor_mentioned": null,
  "next_step": null,
  "objections": ["Security review took three months for their last vendor — IT Security Lead's hesitation"],
  "confidence": "low"
}
```
Note: next_step = null — the rep proposed a CFO follow-up; the prospect said "Maybe — I need to check her calendar, no promises." Not explicitly agreed. The CFO is referenced but not in the speaker list.

TX-005 — Deal-F8767A
```json
{
  "transcript_id": "TX-005",
  "deal_alias": "Deal-F8767A",
  "why_buys": ["Automate service milestones", "Analytics on recognition equity across departments"],
  "pain_points": ["Night-shift teams feel invisible — engagement scores run 20 points lower", "Exec team skeptical after a failed rollout two years ago"],
  "stakeholders": ["Alex Franklin (rep)", "Prospect (HR Director)", "Prospect (People Ops Coordinator)"],
  "budget_signal": "$12k approved under their engagement line",
  "timeline_signal": "Running before their January all-hands",
  "competitor_mentioned": "Nectar (currently mid-pilot with them)",
  "next_step": "Rep presents to the exec team on October 2",
  "objections": ["Exec team skeptical after a failed rollout two years ago", "Would need to beat the Nectar pilot experience"],
  "confidence": "medium"
}
```

TX-006 — Deal-EE195F
```json
{
  "transcript_id": "TX-006",
  "deal_alias": "Deal-EE195F",
  "why_buys": ["Cut the admin time on service awards"],
  "pain_points": ["HR Manager personally spends five hours a month ordering and shipping plaques"],
  "stakeholders": ["Alex Franklin (rep)", "Prospect (HR Manager)"],
  "budget_signal": null,
  "timeline_signal": "Fiscal year resets in February, so a Q1 start is realistic",
  "competitor_mentioned": null,
  "next_step": "Rep sends one-page overview; HR Manager forwards it to the COO this week",
  "objections": ["COO usually prefers building things in-house"],
  "confidence": "medium"
}
```
Note: budget = null — prospect said "Budget isn't the issue — time is," which is a qualitative signal, not a stated amount. Competitor = null — the alternative raised is building internally (an in-house option, not a named vendor); logged under objections instead.

Arithmetic: none required — no totals, rates, or derived figures were asked for or computed. All quantities above (30% turnover, $25k, $40k, $12k, $15k threshold, 5 hours/month, 20-point gap, 12 locations, 6–8 weeks) are quoted verbatim from prospect lines.
next-to-close0.75
Plain-English question: Which 3 deals are closest to signature right now?
Prompt — what the model saw
Which 3 deals are closest to signature right now? Use the CRM data and the Slack thread; Slack often has the real status before the CRM does. For each: deal, amount, why it is close, what is left. Under 250 words.
cost $0.0027 · 18s · in 4,526 / out 732 tok
✓ exactly_three✗ top3_named✓ slack_first✓ blocked_excluded

✓ Correct Answer (ground truth)

Plain-English answer: Deal-547B2B ($11,200), Deal-B7EBD1 ($9,000), and Deal-403845 ($9,000) — all DS5 COMMIT deals closing September 10-11, with what is left on each.
blocked_excludeddeal_alias=Deal-2465CE; reason=champion left + procurement freeze (Slack #sales-sparring 2026-09-04)
exactly_three1
ranking_ruleSlack signature-imminent deal ranks first; remaining ranked by stage (DS5>DS4), forecast (COMMIT>BEST_CASE>PIPELINE), close_date asc, amount desc
top3rank=1; deal_alias=Deal-547B2B; amount=11200; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11 | rank=2; deal_alias=Deal-B7EBD1; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-10 | rank=3; deal_alias=Deal-403845; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11
graded checksexactly_three · top3_named · slack_first · blocked_excluded

Model's actual answer

show the model's full answer
Top 3 closest to signature (Slack threads override CRM where they conflict):

1. Deal-547B2B — $11,200 (Alex Franklin, DS5/COMMIT, close 2026-09-11)
Why close: Slack (09-04 09:12) — redlines returned clean, signing page is out, VP People stated they sign the next day. Signature-imminent.
Left: Countersignature only.

2. Deal-403845 — $9,000 (CRM owner Alex Franklin, DS5/COMMIT, close 2026-09-11)
Why close: Slack (Dana Mercer, 09-04 10:02) — order form is with their finance team; described as "moving fine." Note: Dana reports this deal, but the CRM lists Alex Franklin as owner — flagging the discrepancy, not resolving it.
Left: Finance approval of the order form, then signature.

3. Deal-A2B47C — $6,360 (Alex Franklin, DS5/COMMIT, close 2026-09-11)
Why close: Slack (09-04 09:44) — "still warm," progressing at normal legal-review pace; DS5/COMMIT with a 09-11 close date.
Left: Legal review to conclude, then signature.

Arithmetic: 11,200 + 9,000 + 6,360 = $26,560 combined.

Excluded despite CRM status: Deal-2465CE ($5,400, DS5/COMMIT) — Slack says the champion left and procurement froze new vendors; Dana is pulling it from commit (Q4 realistically). Deal-D348E1 ($13,770) is also warm but its close date (2026-10-15) is later than Deal-A2B47C's (2026-09-11), so it ranks 4th.

No data is missing for these three beyond what's noted.
product-gap-detectionPASS 1.00
Plain-English question: Which prospect complaints are true product gaps versus rollout or enablement issues?
Prompt — what the model saw
Surface true product gaps raised by prospects in active deals. For each candidate: quote the prospect line, classify as real gap, plan gate, or rollout/enablement issue using the product docs excerpt, and name the deal and amount. Only real gaps go in the summary. Rep language does not count as prospect voice.
cost $0.0076 · 68s · in 12,410 / out 2,394 tok
✓ real_gaps✓ plan_gate_classified✓ enablement_classified✓ gap_summary

✓ Correct Answer (ground truth)

Plain-English answer: Two real gaps (SCIM provisioning, ADP Workforce Now integration) separated from plan gates and enablement issues, each with the prospect quote and deal.
allowed_newTG-001, TG-002, TG-003, TG-004, TG-005, SCIM user provisioning, ADP Workforce Now integration, Custom report builder
enablement_issuestranscript_id=TG-004; deal_alias=Deal-84DBA6; amount=16000; feature=Slack integration recognition adoption; quote=We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.; why_enablement=Slack integration exists on all plans; managers were never trained
excluded_rep_voicetranscript_id=TG-005; line=Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap.; reason=rep voice, not prospect-raised
plan_gatestranscript_id=TG-003; deal_alias=Deal-CFE7F4; amount=32000; feature=Custom report builder; tier_required=Enterprise; quote=I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?
real_gapstranscript_id=TG-001; deal_alias=Deal-EC3025; amount=62000; feature=SCIM user provisioning; quote=We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.; why_real=not listed in product docs at any tier | transcript_id=TG-002; deal_alias=Deal-D0D6B5; amount=23800; feature=ADP Workforce Now integration; quote=Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.; why_real=supported HRIS list is Workday/BambooHR/Gusto/Rippling only
graded checksreal_gaps · plan_gate_classified · enablement_classified · gap_summary

Model's actual answer

show the model's full answer
Note on data: transcripts_gaps.csv contains no amount, stage, or pipeline fields. Deal amounts cannot be named — they are not in the provided data. Likewise, "active" status cannot be verified from the file; all deals in the file are treated as in-scope as given.

CANDIDATES (prospect voice only)
─────────────────────────────────

1) TG-001 — Deal-EC3025 — AMOUNT NOT IN DATA
   Prospect (IT Security Lead): "We need SCIM user provisioning — our security team will not approve any tool that requires manual account management."
   Classification: REAL GAP.
   Docs basis: the plans table lists no SCIM capability on any tier, and "Supported integrations" states: "SCIM user provisioning and ADP Workforce Now integrations are NOT currently listed as supported capabilities." Not a tier gate (nothing to upgrade to) and not enablement (feature does not exist).

2) TG-002 — Deal-D0D6B5 — AMOUNT NOT IN DATA
   Prospect (HRIS Manager): "Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us."
   Classification: REAL GAP.
   Docs basis: HRIS integrations are "Workday, BambooHR, Gusto, Rippling" (Pro and above); ADP Workforce Now is explicitly "NOT currently listed as supported." No plan offers it, so it is not a plan gate.

3) TG-003 — Deal-CFE7F4 — AMOUNT NOT IN DATA
   Prospect (VP People): "I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?"
   Classification: PLAN GATE (not a real gap).
   Docs basis: "Custom report builder" is listed in the capability table — Enterprise tier only. The capability exists; the prospect simply isn't on the tier that includes it. (Rep line confirming Enterprise tier is excluded as evidence per instruction, but the docs table is the classification basis regardless.)

4) TG-004 — Deal-84DBA6 — AMOUNT NOT IN DATA
   Prospect (People Ops Manager): "We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it."
   Classification: ROLLOUT/ENABLEMENT ISSUE (not a real gap).
   Docs basis: "Slack and Microsoft Teams integration" is available on all plans and the integration is already turned on; the prospect names lack of training ("nobody ever trained them on it") as the cause. No missing capability is claimed.

5) TG-005 — Deal-36C33F — EXCLUDED
   The only gap claim ("we don't have a native mobile app for hourly workers yet") is spoken by Alex Franklin (rep). Per the rule that rep language does not count as prospect voice, this is not a prospect-raised gap. The prospect line in this transcript ("Good to know. The web version should be fine for our office staff for now.") raises no gap.

ARITHMETIC
──────────
Transcripts reviewed: 5
Prospect-raised gap candidates (excludes rep-only lines): 4 (TG-001..TG-004; TG-005 excluded)
  - Real gaps: 2 (TG-001, TG-002)
  - Plan gates: 1 (TG-003)
  - Rollout/enablement: 1 (TG-004)
Check: 2 + 1 + 1 = 4. ✓

SUMMARY — REAL GAPS ONLY
────────────────────────
1. SCIM user provisioning — absent from all plans and from the supported-integrations list (product docs snapshot 2026-08). Deal-EC3025. Amount: not provided in the data.
2. ADP Workforce Now integration — absent from the HRIS list (Workday, BambooHR, Gusto, Rippling only) and explicitly not listed as supported. Deal-D0D6B5, stated as a dealbreaker by the prospect. Amount: not provided in the data.

Both real gaps share one docs statement: "SCIM user provisioning and ADP Workforce Now integrations are NOT currently listed as supported capabilities."
rep-performance 5 tests
stale-pipeline-by-repPASS 1.00
Plain-English question: Which open deals have had no email, call, or meeting in the last 7 days?
Prompt — what the model saw
The snapshot date is 2026-09-05. Using the open deals and the engagements-by-deal table, list every open deal with no logged email, call, or meeting in the last 7 days. Compute recency from the engagements table's last_email/last_call/last_meeting fields, NOT from any date field on the deal (the deal's last_contacted_field is unreliable). For each stale deal give: deal alias, owner name, stage, amount, days since last contact. Group by owner, order by amount descending within each owner. Then state, per owner, how many stale deals they carry and the total stale amount. (Do not write outreach text.)
cost $0.0664 · 441s · in 79,738 / out 35,199 tok
✓ stale_recall_ge_80✓ grouped_by_owner✓ jake_top_amount✓ kevin_stale_amt✓ jenn_stale_amt✓ stale_count

✓ Correct Answer (ground truth)

Plain-English answer: 65 open deals are stale; the largest is Deal-2D1F1B at $240,000, with Bryce Harmon carrying the most stale dollars at $626,243.
The answer in one line: jake_amt = 626,243; jenn_amt = 261,645; kevin_amt = 252,905.03; top_amount = 240,000; stale_count = 65; top_alias = Deal-2D1F1B
jake_amt626,243
jenn_amt261,645
kevin_amt252,905.03
stale_count65
top_amount240,000
top_aliasDeal-2D1F1B
graded checksstale_recall_ge_80 · grouped_by_owner · jake_top_amount · kevin_stale_amt · jenn_stale_amt · stale_count

Model's actual answer

show the model's full answer
STALE OPEN DEALS — no email/call/meeting logged in the 7 days before 2026-09-05

METHOD
- Recency per deal = max(last_email, last_call, last_meeting) from engagements_by_deal_90d.csv. The deal-level last_contacted_field was NOT used.
- Days since = 2026-09-05 minus that date. Stale = more than 7 days (last contact on or before 2026-08-28). Note: no deal's last contact falls exactly on 2026-08-29, so the choice of 7-day boundary does not change the list.
- Owner groups ordered by total stale amount (descending); within each owner, deals ordered by amount descending.

================================================================
BRYCE HARMON — 13 stale deals, $626,243.00 total
================================================================
Deal alias | stage | amount | days since last contact
Deal-2D1F1B | DS1 | $240,000.00 | 81 (last_meeting 2026-06-16)
Deal-66D1FC | DS1 | $99,000.00 | 16 (last_email 2026-08-20)
Deal-950043 | DS1 | $70,000.00 | 19 (last_email 2026-08-17)
Deal-B23205 | DS1 | $45,000.00 | 16 (last_email / last_meeting 2026-08-20)
Deal-7BBDFA | DS3 | $37,440.00 | 46 (last_email 2026-07-21)
Deal-332637 | DS2 | $36,000.00 | 9 (last_email 2026-08-27)
Deal-1BEEBF | DS1 | $31,500.00 | 19 (last_email 2026-08-17)
Deal-C5658B | DS1 | $23,400.00 | 16 (last_email 2026-08-20)
Deal-40522D | DS3 | $21,000.00 | 19 (last_email 2026-08-17)
Deal-F0EBBB | DS3 | $11,400.00 | 24 (last_email 2026-08-12)
Deal-E25A09 | DS1 | $6,000.00 | 9 (last_email 2026-08-27)
Deal-C9C286 | DS2 | $5,502.00 | 9 (last_email 2026-08-27)
Deal-012CB1 | DS1 | $1.00 | 23 (last_email 2026-08-13)

Arithmetic: 240,000 + 99,000 + 70,000 + 45,000 + 37,440 + 36,000 + 31,500 + 23,400 + 21,000 + 11,400 + 6,000 + 5,502 + 1 = $626,243.00

================================================================
DANA MERCER — 14 stale deals, $261,645.00 total
================================================================
Deal-44EA29 | DS2 | $60,000.00 | 10 (last_email 2026-08-26)
Deal-E51FB7 | DS2 | $43,875.00 | 12 (last_call 2026-08-24)
Deal-B42F46 | DS1 | $27,000.00 | 19 (last_email 2026-08-17)
Deal-BA3DDC | DS3 | $23,400.00 | 15 (last_call 2026-08-21)
Deal-9DDE86 | DS2 | $20,000.00 | 15 (last_email 2026-08-21)
Deal-215CCA | DS3 | $18,900.00 | 17 (last_meeting 2026-08-19)
Deal-5EED42 | DS3 | $16,250.00 | 11 (last_email / last_call 2026-08-25)
Deal-57887A | DS2 | $15,000.00 | 8 (last_email 2026-08-28)
Deal-B7EBD1 | DS5 | $9,000.00 | 16 (last_email 2026-08-20)
Deal-3974EB | DS4 | $9,000.00 | 8 (last_email / last_meeting 2026-08-28)
Deal-F40F04 | DS2 | $8,100.00 | 15 (last_email / last_meeting 2026-08-21)
Deal-87DDD1 | DS1 | $5,000.00 | 19 (last_email 2026-08-17)
Deal-F336B6 | DS3 | $4,200.00 | 15 (last_email 2026-08-21)
Deal-0660B4 | DS4 | $1,920.00 | 16 (last_meeting 2026-08-20)

Arithmetic: 60,000 + 43,875 + 27,000 + 23,400 + 20,000 + 18,900 + 16,250 + 15,000 + 9,000 + 9,000 + 8,100 + 5,000 + 4,200 + 1,920 = $261,645.00

================================================================
COLE INGRAM — 18 stale deals, $252,905.03 total
================================================================
Deal-D04904 | DS2 | $58,529.25 | 11 (last_email 2026-08-25)
Deal-B25F40 | DS3 | $40,000.00 | 8 (last_email 2026-08-28)
Deal-813836 | DS2 | $32,175.00 | 11 (last_email 2026-08-25)
Deal-1BA595 | DS2 | $31,750.00 | 11 (last_email 2026-08-25)
Deal-CFE1E8 | DS3 | $18,000.00 | 11 (last_email 2026-08-25)
Deal-CD47A6 | DS2 | $12,168.00 | 11 (last_email 2026-08-25)
Deal-627646 | DS3 | $11,193.00 | 11 (last_email 2026-08-25)
Deal-FF809F | DS2 | $7,781.20 | 11 (last_email 2026-08-25)
Deal-AF932D | DS2 | $7,225.40 | 11 (last_email 2026-08-25)
Deal-A71728 | DS2 | $6,947.50 | 11 (last_email 2026-08-25)
Deal-8BC9F5 | DS2 | $5,616.00 | 10 (last_email 2026-08-26)
Deal-175395 | DS3 | $4,779.88 | 11 (last_email 2026-08-25)
Deal-481E24 | DS3 | $4,140.00 | 10 (last_call 2026-08-26)
Deal-C7F9BF | DS2 | $3,360.00 | 11 (last_email 2026-08-25)
Deal-2F3A66 | DS3 | $3,334.80 | 11 (last_email 2026-08-25)
Deal-342E96 | DS2 | $2,700.00 | 24 (last_email 2026-08-12)
Deal-E568D5 | DS3 | $1,875.00 | 11 (last_email 2026-08-25)
Deal-FD9F4E | DS5 | $1,330.00 | 10 (last_email 2026-08-26)

Arithmetic: 58,529.25 + 40,000 + 32,175 + 31,750 + 18,000 + 12,168 + 11,193 + 7,781.20 + 7,225.40 + 6,947.50 + 5,616 + 4,779.88 + 4,140 + 3,360 + 3,334.80 + 2,700 + 1,875 + 1,330 = $252,905.03

================================================================
ALEX FRANKLIN — 18 stale deals, $102,336.00 total
================================================================
Deal-CC08D1 | DS1 | $24,000.00 | 16 (last_email 2026-08-20)
Deal-E73427 | DS3 | $18,000.00 | 10 (last_email / last_meeting 2026-08-26)
Deal-885F45 | DS2 | $9,300.00 | 12 (last_email 2026-08-24)
Deal-C2FF3C | DS1 | $8,316.00 | 10 (last_email 2026-08-26)
Deal-0D2F7A | DS3 | $5,100.00 | 12 (last_call 2026-08-24)
Deal-6C60D4 | DS3 | $4,800.00 | 12 (last_call 2026-08-24)
Deal-13FEBD | DS2 | $4,680.00 | 12 (last_call 2026-08-24)
Deal-9D0060 | DS3 | $3,840.00 | 12 (last_email 2026-08-24)
Deal-690476 | DS2 | $3,600.00 | 18 (last_call 2026-08-18)
Deal-C6D97A | DS4 | $3,240.00 | 8 (last_email 2026-08-28)
Deal-EE195F | DS3 | $3,120.00 | 8 (last_email 2026-08-28)
Deal-278DEC | DS3 | $2,700.00 | 8 (last_email 2026-08-28)
Deal-635B8E | DS3 | $2,600.00 | 18 (last_email 2026-08-18)
Deal-6883F3 | DS1 | $2,400.00 | 16 (last_email / last_meeting 2026-08-20)
Deal-4A13AD | DS3 | $2,160.00 | 26 (last_email 2026-08-10)
Deal-F67D31 | DS2 | $1,800.00 | 8 (last_email 2026-08-28)
Deal-5FDCE4 | DS3 | $1,600.00 | 12 (last_email 2026-08-24)
Deal-BA571A | DS4 | $1,080.00 | 18 (last_email 2026-08-18)

Arithmetic: 24,000 + 18,000 + 9,300 + 8,316 + 5,100 + 4,800 + 4,680 + 3,840 + 3,600 + 3,240 + 3,120 + 2,700 + 2,600 + 2,400 + 2,160 + 1,800 + 1,600 + 1,080 = $102,336.00

================================================================
PER-OWNER SUMMARY
================================================================
Bryce Harmon   | 13 deals | $626,243.00
Dana Mercer    | 14 deals | $261,645.00
Cole Ingram    | 18 deals | $252,905.03
Alex Franklin  | 18 deals | $102,336.00
TOTAL          | 63 deals | $1,243,129.03 (626,243 + 261,645 + 252,905.03 + 102,336)

================================================================
DATA GAPS AND CAVEATS (stated explicitly, not invented around)
================================================================
1. Two open deals have NO row in engagements_by_deal_90d.csv, so recency cannot be computed from that table and days-since is undefined. They are excluded from the list and totals above:
   - Deal-3EED2C | owner Alex Franklin | DS2 | $7,200 (deal 64623982954)
   - Deal-57FF13 | owner Elena Sinclair | DS1 | $2,100 (deal 64524667574)
   Both were created 2026-09-01/09-03 and have no logged email/call/meeting in the engagements table at all. If treated as stale (no contact logged in last 7 days), the counts become: Alex Franklin 19 deals / $109,536.00; Elena Sinclair 1 deal / $2,100.00; overall 65 deals / $1,252,429.03.

2. Future-dated last_meeting values exist in the engagements table (e.g., 2026-09-09 through 2026-10-02, all after the 2026-09-05 snapshot). Taken at face value per instructions, these keep 10 deals out of the stale list. If future-dated meetings are excluded from recency, these 10 deals would also be stale (days-since computed from the last past-dated engagement):
   - Deal-944310 (Dana Mercer) 33d; Deal-7599B8 (Dana Mercer) 18d; Deal-3795AD (Bryce Harmon) 8d; Deal-01E193 (Bryce Harmon) 8d; Deal-C1FA6D (Bryce Harmon) 16d; Deal-927338 (Bryce Harmon) 18d; Deal-A414F6 (Bryce Harmon) 19d; Deal-819506 (Alex Franklin) 8d; Deal-8BA24E (Farid Osman) 8d; Deal-8FDCD2 (Farid Osman) 15d.
   The base answer above uses the fields as given.
activity-mix-vs-outcomePASS 1.00
Plain-English question: Per rep, what did the last 30 days of emails, calls, and meetings look like next to the deals that advanced?
Prompt — what the model saw
The snapshot date is 2026-09-05. For each sales rep (owner), compute the last-30-day totals of emails, calls, and meetings from the engagements table, and the number of deals that entered DS2 in the last 30 days (t_ds2 within 30 days of the snapshot). Report per rep: the activity mix as percentages (emails/calls/meetings share of that rep's total activities), total activities per DS2 entry, and rank the reps by that efficiency ratio (lowest activities per DS2 entry = most efficient). Name the most efficient rep and the highest-volume rep, and note if they differ. Do not attribute cause.
cost $0.0418 · 436s · in 31,987 / out 31,499 tok
✓ alex_ds2_30d✓ justin_volume_high✓ efficient_named✓ mix_present✓ no_causal_claim

✓ Correct Answer (ground truth)

Plain-English answer: Alex Franklin is the standout signal with 18 deals entering DS2 in 30 days (next closest: Bryce Harmon with 4), alongside per-rep activity totals.
The answer in one line: alex_ds2_30d = 18; bryce_ds2_30d = 4; note = legacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
alex_ds2_30d18
bryce_ds2_30d4
notelegacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
graded checksalex_ds2_30d · justin_volume_high · efficient_named · mix_present · no_causal_claim

Model's actual answer

show the model's full answer
Done. Computed directly from the three provided files (window: t_ds2 in 2026-08-06..2026-09-05, i.e. within 30 days of the 2026-09-05 snapshot, inclusive at both ends; activity = sum of emails_30d/calls_30d/meetings_30d across each owner's deals).

PER-REP RESULTS

Alex Franklin (84342457) — 67 deals in deals_open
  Activities: emails 307 + calls 36 + meetings 41 = 384
  Mix: 307/384 = 79.9% emails, 36/384 = 9.4% calls, 41/384 = 10.7% meetings
  DS2 entries in window (18): Deal-403845, Deal-1FC049, Deal-3EED2C, Deal-7FA0C3, Deal-E531A6, Deal-5296C9, Deal-36C33F, Deal-EE195F, Deal-F436DA, Deal-317E6F, Deal-D1E6C2, Deal-D9A72E, Deal-CA5E44, Deal-4F775F, Deal-898FC5, Deal-46988D, Deal-E73427, Deal-92D97D
  Efficiency: 384 / 18 = 21.33 activities per DS2 entry

Bryce Harmon (119337721) — 35 deals
  Activities: emails 162 + calls 0 + meetings 43 = 205
  Mix: 162/205 = 79.0% emails, 0/205 = 0.0% calls, 43/205 = 21.0% meetings
  DS2 entries (4): Deal-25F752 (2026-08-10), Deal-D73B89 (2026-09-03), Deal-CA7DC0 (2026-08-12), Deal-1CCE5C (2026-08-06)
  Efficiency: 205 / 4 = 51.25

Cole Ingram (83155924) — 22 deals
  Activities: emails 96 + calls 14 + meetings 1 = 111
  Mix: 96/111 = 86.5% emails, 14/111 = 12.6% calls, 1/111 = 0.9% meetings
  DS2 entries (2): Deal-42326B (2026-08-26), Deal-1BA595 (2026-08-12)
  Efficiency: 111 / 2 = 55.50

Dana Mercer (83155923) — 24 deals
  Activities: emails 84 + calls 18 + meetings 11 = 113
  Mix: 84/113 = 74.3% emails, 18/113 = 15.9% calls, 11/113 = 9.7% meetings
  DS2 entries (1): Deal-57887A (2026-08-07)
  Efficiency: 113 / 1 = 113.00

Farid Osman (716654662) — 7 deals
  Activities: emails 38 + calls 0 + meetings 34 = 72
  Mix: 38/72 = 52.8% emails, 0/72 = 0.0% calls, 34/72 = 47.2% meetings
  DS2 entries (1): Deal-499BF6 (2026-08-26)
  Efficiency: 72 / 1 = 72.00

Elena Sinclair (701163055) — 1 deal (Deal-57FF13)
  Activities: 0 (no row in engagements table for Deal-57FF13)
  Mix: not computable (0 activities)
  DS2 entries (0): ratio N/A — excluded from ranking

RANKING BY ACTIVITIES PER DS2 ENTRY (lowest = most efficient; reps with >=1 DS2 entry)
  1. Alex Franklin — 21.33
  2. Bryce Harmon — 51.25
  3. Cole Ingram — 55.50
  4. Farid Osman — 72.00
  5. Dana Mercer — 113.00

HIGHEST-VOLUME REPS (total activities)
  1. Alex Franklin — 384; 2. Bryce Harmon — 205; 3. Dana Mercer — 113; 4. Cole Ingram — 111; 5. Farid Osman — 72; 6. Elena Sinclair — 0

MOST EFFICIENT: Alex Franklin (21.33 activities per DS2 entry)
HIGHEST VOLUME: Alex Franklin (384 activities)
They do not differ — the same rep holds both.

DATA NOTES (not interpretation)
- Deal-3EED2C (Alex Franklin) and Deal-57FF13 (Elena Sinclair) have no row in the engagements table; their activity is absent from the totals above.
- Deal-EE195F, Deal-D9A72E (Alex Franklin) and Deal-1CCE5C (Bryce Harmon) have t_ds2 = 2026-08-06, exactly 30 days before the snapshot; counted as within the window. Excluding the boundary would drop those reps' counts to 16 and 3 respectively (ratios 24.00 and 68.33) and reorder ranks 2-4.
- Gavin Porter and Hugo Lindqvist are archived with no deals in deals_open; no activity or DS2 data exists for them.
ae-qtd-scorecardPASS 1.00
Plain-English question: Build Alex Franklin's QTD scorecard: bookings versus quota, pipeline, and conversion.
Prompt — what the model saw
Build a QTD scorecard for the named AE as of 2026-09-05: bookings vs quota with attainment percent, new vs expansion split, active pipeline by stage with amount, rolling 90-day DS2-to-won rate, win and loss counts with top loss reason, activity volume by type in the last 30 days. Exclude any closed-won deal dated before the quarter. Close with three coaching observations grounded in the numbers.
cost $0.0276 · 181s · in 30,733 / out 16,027 tok
✓ bookings✓ quota✓ attainment✓ new_expansion_split✓ top_loss_reason✓ coaching_obs

✓ Correct Answer (ground truth)

Plain-English answer: 75% attainment ($150,000 on a $200,000 quota), $113,500 new versus $36,500 expansion, pipeline by stage, and the rolling DS2-to-won rate.
The answer in one line: quota = 200000; bookings = 150000; new_bookings = 113500; expansion_bookings = 36500; open_deals_count = 125; attainment_pct = 75; owner = Alex Franklin (+12 more below)
quota200,000
attainment_pct75
bookings150,000
expansion_bookings36,500
expansion_count3
losses_qtd27
new_bookings113,500
new_count5
open_deals_count125
wins_qtd8
activity_30d_totalsemails_30d=807; calls_30d=112; meetings_30d=128; notes_30d=50
allowed_new18 items — e.g. Deal-A1C3E5, Deal-B7D2F4, Deal-C9E1A6, Deal-D4B8C2
ds2_to_wonwindow_start=2026-06-07; won_in_window=8; closed_in_window=35; rate=0.2286
ownerAlex Franklin
pipeline_by_stage{"DS1": {"count": 20, "amount": 284621}, "DS2": {"count": 28, "amount": 353760}, "DS3": {"count": 67, "amount": 552705}, "DS4": {"count": 5, "amount": 23574}, "DS5": {"count": 5, "amount": 45730}}
prior_quarter_excludeddeal_alias=Deal-B3E6F1; close_date=2026-06-20; amount=24000
quarter2026-Q3
snapshot_date2026-09-05
top_loss_reasonreason=Lost- Timing (1 year or more); count=13
graded checksbookings · quota · attainment · new_expansion_split · top_loss_reason · coaching_obs

Model's actual answer

show the model's full answer
AE QTD SCORECARD — Alex Franklin, 2026-Q3, as of 2026-09-05 (quota.csv snapshot 2026-09-05, quota $200,000)

BOOKINGS vs QUOTA
QTD closed-won (close_date 2026-07-01..2026-09-30): 8 deals, $150,000
  Excluded: Deal-B3E6F1 ($24,000, close 2026-06-20) — closed-won before quarter start
Attainment: $150,000 / $200,000 = 75.0%
Remaining to quota: $50,000

NEW vs EXPANSION (of the $150,000)
  New:      5 deals — $40,000 + $35,000 + $21,000 + $11,000 + $6,500 = $113,500 (75.7%)
  Expansion: 3 deals — $20,000 + $12,000 + $4,500 = $36,500 (24.3%)
  Check: $113,500 + $36,500 = $150,000 ✓

ACTIVE PIPELINE BY STAGE (all open deals, n=125)
  DS1:  20 deals   $284,621
  DS2:  28 deals   $353,760
  DS3:  67 deals   $552,705
  DS4:   5 deals    $23,574
  DS5:   5 deals    $45,730
  Total: 125 deals  $1,260,390
  (Subset closing on/before 2026-09-30: 22 deals, $109,363)

ROLLING 90-DAY DS2-TO-WON RATE (entered_ds2 2026-06-07..2026-09-05)
  Deals that entered DS2 in window: 111 — 8 won, 27 lost, 76 still open
  Won / entered: 8 / 111 = 7.2%
  Won / closed outcomes: 8 / (8+27) = 8/35 = 22.9%

WINS AND LOSSES QTD
  Wins: 8 ($150,000, avg $18,750)
  Losses: 27 (avg $12,195)
  Loss reasons (count / amount):
    Lost- Timing (1 year or more) — 13 / $184,681  ← top loss reason
    MIA — 5 / $45,831
    Competitor — 5 / $49,020
    Lost DM — 2 / $17,940
    Feature Request — 1 / $21,000
    Lost- Does not fit ICP (write in notes) — 1 / $10,800

ACTIVITY, LAST 30 DAYS (ae_engagements.csv, summed over 161 deal rows)
  Emails: 807
  Calls: 112
  Meetings: 128
  Notes: 50
  Total: 1,097
  (On the 9 won-deal rows incl. pre-quarter Deal-B3E6F1: 99 emails / 33 calls / 25 meetings / 24 notes)

COACHING OBSERVATIONS
1. Attainment is 75.0% ($150,000 of $200,000) with 25 days left in-quarter, but only $109,363 of open pipeline closes on/before 9/30 — coverage of the $50,000 gap is ~2.2x in-quarter dollars and depends on DS5 (Deal-403845, Deal-547B2B, Deal-A2B47C, $26,560 in-quarter) converting late-stage on time. Deal-547B2B ($11,200) has been in-cycle since entered_ds2 2026-04-10 — a 5-month-old late-stage deal carrying 22% of in-quarter-close DS5 dollars.
2. Timing is the dominant failure mode: 13 of 27 QTD losses ($184,681, 48.4% of loss dollars) are "Lost- Timing (1 year or more)", and the rolling-90d DS2-to-won rate is 7.2% (8 of 111 entries; 76 still open), so qualification — not just closing — is where deals are dying. Loss dollars per deal ($12,195 avg) are running near won-deal size ($18,750 avg), so these aren't unqualified junk; they're deals getting to DS2 and stalling.
3. Activity is heavily email-weighted: 807 emails vs 112 calls and 128 meetings (7.2 emails per call), and won deals show the live-interaction pattern (on the 9 won rows: 33 calls + 25 meetings + 24 notes vs 99 emails). The 28 open DS2 deals ($353,760 — the largest early-stage dollar pool after DS3) are where call/meeting coverage would most directly move the 7.2% conversion.
multithreading-gapPASS 1.00
Plain-English question: Which open deals are single-threaded or under-threaded?
Prompt — what the model saw
Find every open deal that is single-threaded (fewer than 2 active contacts) or under-threaded (fewer than 3, or all contacts in one persona). Active means engaged in the last 60 days and not marked former. Personas: economic buyer, champion, HR admin, IT security, finance. For each flagged deal: amount, stage, active contact count, personas present, personas missing, the most valuable persona to add given the stage, and any on-file unengaged contact at that company who fits; otherwise say none on file.
cost $0.0087 · 118s · in 5,347 / out 7,212 tok
✓ flagged_recall✓ persona_language✓ single_thread_term✓ on_file_fit

✓ Correct Answer (ground truth)

Plain-English answer: Deals with fewer than 2 active contacts (single-threaded) or fewer than 3, or all in one persona (under-threaded), each flagged with the missing persona.
The answer in one line: active_window_start = 2026-07-07
active_window_start2026-07-07
allowed_new59 items — e.g. C-077A0E, C-0D15DF, C-1BDFE0, C-2AE3AA
best_persona_ruleDS2/DS3: economic buyer > IT security > finance > HR admin > champion; DS4/DS5: economic buyer > finance > IT security > HR admin > champion; first persona not present among active contacts
flag_ruleflag if active_contacts < 2 (single-threaded) or active_contacts < 3 or all active contacts share one persona (under-threaded); active = engaged within 60 days and not former
flaggeddeal_alias=Deal-36C33F; amount=15000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-50D386; amount=36000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5408B0; amount=14850; stage=DS4; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5BFE3B; amount=18000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-885F45; amount=9300; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=IT security | deal_alias=Deal-92D97D; amount=60000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer … (+5 more)
healthy_not_flaggedDeal-4B0BEB, Deal-84DBA6, Deal-D348E1
persona_universeeconomic buyer, champion, HR admin, IT security, finance
graded checksflagged_recall · persona_language · single_thread_term · on_file_fit

Model's actual answer

show the model's full answer
DATA GAPS (stated up front): the provided files contain no deal amount, no stage, and no open/closed flag. I cannot report amount or stage for any deal, cannot condition "most valuable persona" on stage, and cannot verify which deals are open — I analyze every deal present in deal_contacts.csv and flag those that are under-threaded. Below, "most valuable persona to add" therefore uses a stage-independent default priority: economic buyer > champion > HR admin > finance > IT security (with stage data, late-stage deals would instead prioritize finance / IT security).

ACTIVE-CONTACT RULE (arithmetic): as of 2026-09-24, the 60-day window opens 2026-07-26 (24 days to 2026-08-31... = 24 + 31 + 5 = 60 days back to 2026-07-26). Active = last_engaged_date >= 2026-07-26 AND is_former = false. (A 2026-09-23 as-of date gives cutoff 2026-07-25 — no contact falls between, so classifications are identical.)

Excluded as inactive: CT-F2C1AE (former), CT-405B45 (former), CT-86B22F (former), CT-A902AE (2026-06-01, 115 days), CT-913581 (2026-06-20, 96 days).

11 of 14 deals are flagged. Not flagged: Deal-4B0BEB (4 active, 4 personas), Deal-84DBA6 (3 active, 3 personas), Deal-D348E1 (5 active, 5 personas).

FLAGGED DEALS
--------------

1) Deal-EC3025 (61032318100, C-FDD0C7)
   Amount / stage: not in data.
   Active contacts: 1 of 2 (CT-F2C1AE is former) -> single-threaded.
   Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
   Most valuable add: economic buyer (deal's only economic buyer, CT-F2C1AE, is former).
   On-file unengaged fit: CT-6827DB, Chief People Officer, economic buyer.

2) Deal-92D97D (59728118877, C-E23238)
   Amount / stage: not in data.
   Active contacts: 1 of 2 (CT-A902AE stale since 2026-06-01) -> single-threaded.
   Personas present: HR admin. Missing: economic buyer, champion, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: none on file. (Also worth noting: champion CT-A902AE is stale-but-not-former and could be re-engaged.)

3) Deal-50D386 (61055128146, C-EB10E4)
   Amount / stage: not in data.
   Active contacts: 2 of 2 -> under-threaded (2 < 3).
   Personas present: champion, HR admin. Missing: economic buyer, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: CT-A1C4B3, Chief People Officer, economic buyer.

4) Deal-D0D6B5 (60081655042, C-32918E)
   Amount / stage: not in data.
   Active contacts: 3 of 3 -> under-threaded (all 3 champions, one persona).
   Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: CT-1FA4DB, Chief People Officer, economic buyer.

5) Deal-5BFE3B (51674270311, C-535D36)
   Amount / stage: not in data.
   Active contacts: 2 of 2 -> under-threaded (2 < 3, and both champions).
   Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: none on file.

6) Deal-36C33F (63739413805, C-077A0E)
   Amount / stage: not in data.
   Active contacts: 1 of 3 (CT-405B45 and CT-86B22F both former) -> single-threaded.
   Personas present: IT security. Missing: economic buyer, champion, HR admin, finance.
   Most valuable add: economic buyer (the deal lost its economic buyer, CT-86B22F, to a former-marked exit).
   On-file unengaged fit: CT-1DB73E, Chief People Officer, economic buyer.

7) Deal-885F45 (60686135564, C-5E8EFB)
   Amount / stage: not in data.
   Active contacts: 2 of 2 -> under-threaded (2 < 3).
   Personas present: economic buyer, champion. Missing: HR admin, IT security, finance.
   Most valuable add: HR admin (top of the default priority among missing; with stage data, late-stage would favor finance / IT security).
   On-file unengaged fit: none on file for HR admin. (CT-B3F25D, IT Security Lead, is on file unengaged at C-5E8EFB but does not fit the recommended persona.)

8) Deal-FCBE5B (62639586615, C-737030)
   Amount / stage: not in data.
   Active contacts: 1 of 1 -> single-threaded.
   Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: none on file.

9) Deal-5408B0 (60182332309, C-2AE3AA)
   Amount / stage: not in data.
   Active contacts: 2 of 2 -> under-threaded (2 < 3).
   Personas present: champion, HR admin. Missing: economic buyer, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: CT-07FA76, Chief People Officer, economic buyer.

10) Deal-C6D97A (62121783047, C-5A8FC2)
    Amount / stage: not in data.
    Active contacts: 3 of 3 -> under-threaded (all 3 champions, one persona).
    Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
    Most valuable add: economic buyer.
    On-file unengaged fit: none on file.

11) Deal-F9A08A (49757401138, C-0D15DF)
    Amount / stage: not in data.
    Active contacts: 1 of 2 (CT-913581 stale since 2026-06-20) -> single-threaded.
    Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
    Most valuable add: economic buyer.
    On-file unengaged fit: CT-697541, Chief People Officer, economic buyer. (Deal's own economic buyer CT-913581 is stale-but-not-former and is also a re-engagement candidate.)

SUMMARY COUNTS
Single-threaded (< 2 active): 5 — Deal-EC3025, Deal-92D97D, Deal-36C33F, Deal-FCBE5B, Deal-F9A08A.
Under-threaded (2 active, or >=3 all one persona): 6 — Deal-50D386, Deal-5408B0, Deal-885F45 (2 active); Deal-D0D6B5, Deal-5BFE3B (2 active, one persona), Deal-C6D97A (3 active, one persona).
Persona recommendations covered by an on-file unengaged contact: 6 of 11 (EC3025, 50D386, D0D6B5, 36C33F, 5408B0, F9A08A). None on file for the recommended persona: 5 (92D97D, 5BFE3B, 885F45, FCBE5B, C6D97A).
talk-track-from-transcriptsPASS 1.00
Plain-English question: From this rep's last 10 calls: how do they open, handle objections, and how often do they lock a next step?
Prompt — what the model saw
From this rep's last 10 calls: what they lead with in the first five minutes, how they handle the three most common objections, how often a concrete next step is agreed (rate), and every competitor a prospect raised. Quote at most one line per finding and only lines that appear in the transcripts. End with two coaching notes.
cost $0.0044 · 40s · in 6,137 / out 1,975 tok
✓ objection_count✓ next_step_rate✓ competitors✓ coaching✓ quote_used

✓ Correct Answer (ground truth)

Plain-English answer: 8 of 10 calls open with the same customer-win story; objection handling, next-step agreement rate, and every competitor raised, with quotes.
The answer in one line: n_transcripts = 10; objection_count = 3; rep = Alex Franklin (+3 more below)
n_transcripts10
objection_count3
allowed_new13 items — e.g. TT-001, TT-002, TT-003, TT-004
competitor_quotesAwardco=We're also in late talks with Awardco — their rewards catalog looks bigger than yours.; Kudos=How are you different from Kudos? Our CEO used them at her last company.
competitors_prospect_raisedAwardco, Kudos
next_stepagreed=7; total=10; rate=0.7
objectionsobjection=budget; count=4; quote=Honestly, budget is locked until next fiscal year — I can't add a new line item right now.; transcript_id=TT-001 | objection=timing; count=3; quote=This is good, but can we revisit it next quarter? Open enrollment starts in October and we're underwater.; transcript_id=TT-002 | objection=in-house; count=3; quote=We already do recognition with a spreadsheet and quarterly gift cards — why would we change?; transcript_id=TT-004
opener_themetheme=opens with a customer recognition win story (400-person retailer, 18% turnover reduction); count=8; total=10; quote=Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it.; transcript_id=TT-001
repAlex Franklin
rep_raised_competitors_excludedWorkhuman
graded checksobjection_count · next_step_rate · competitors · coaching · quote_used

Model's actual answer

show the model's full answer
REP CALL REVIEW — Alex Franklin, last 10 calls (TT-001 … TT-010)

1) WHAT HE LEADS WITH (minute 0)
8 of 10 calls open with the same customer-proof story (400-person retailer, 18% turnover cut). The other two open on logistics/pricing: TT-004 (Deal-403845) and TT-009 (Deal-1E2498). Arithmetic: 8/10 = 80% proof-story opener.
Quote (TT-001, Deal-D348E1, m0): "a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards"

2) THREE MOST COMMON OBJECTIONS AND HOW HE HANDLES THEM

A. Budget (5 of 10: TT-001 Deal-D348E1, TT-003 Deal-547B2B, TT-006 Deal-60C2C2, TT-010 Deal-84DBA6 at m6; TT-004 Deal-403845 at m11).
Handle: in 4 of 5 he reframes cost as funded by turnover savings, citing the $210k backfill figure. On the pure "committee" version he parks the deal ("Understood — I'll leave it with you" — TT-004, Deal-403845, m12).
Quote (TT-003, Deal-547B2B, m8): "that retailer saved about $210k in avoided backfills, which is how their finance team signed off."

B. Timing / no urgency (4 of 10: TT-002 Deal-5408B0, TT-005 Deal-C61CF7, TT-008 Deal-D9A12F at m6 "revisit it next quarter"; TT-007 Deal-EDC141 at m14 "no urgency").
Handle: 3 of 4 get the identical 90-day pilot counter; the no-urgency one is conceded ("Fair enough." — TT-007, Deal-EDC141, m15).
Quote (TT-002, Deal-5408B0, m8): "What if we scope a 90-day pilot with one department so you have internal data before next quarter's planning?"

C. Status quo / spreadsheets (3 of 10: TT-004 Deal-403845, TT-007 Deal-EDC141, TT-009 Deal-1E2498, all m6).
Handle: 3 of 3 with the identical automation + analytics answer.
Quote (TT-009, Deal-1E2498, m8): "the difference is automation: milestones fire without HR lifting a finger, and you get analytics on who is being recognized."

3) CONCRETE NEXT STEP AGREED — RATE
Next step proposed in 7/10 calls (TT-001, 002, 003, 005, 006, 008, 009) and accepted in 7/7. Not proposed in TT-004, TT-007, TT-010 (all three ended with an unresolved objection).
- Rate on calls where asked: 7/7 = 100%
- Rate across all 10 calls: 7/10 = 70%
Quote (TT-005, Deal-C61CF7, m15): "Yes, Thursday at 2pm works — send the invite and I'll bring our HRIS manager."

4) COMPETITORS RAISED BY PROSPECTS
- Awardco — TT-003, Deal-547B2B. Quote (m4): "We're also in late talks with Awardco — their rewards catalog looks bigger than yours."
- Kudos — TT-007, Deal-EDC141. Quote (m4): "How are you different from Kudos? Our CEO used them at her last company."
Note: Workhuman (TT-005, Deal-C61CF7, m2) was raised by the rep, not the prospect, so it is excluded from this list.

TWO COACHING NOTES
1. The openers, objection answers, and close line are verbatim-identical across calls. The one rep that departs from the script (TT-004, TT-009) and the "committee/no urgency" variants are exactly where no next step gets booked — he has no second-close or re-engagement motion when the standard ask fails. Build a fallback ask (e.g., a 15-min data share or send-the-deck-plus-check-in date) instead of "Understood — I'll leave it with you."
2. When a competitor surfaces (Awardco, Kudos), he answers once at m5 and never returns to it — and in TT-003 the deal still gets a next step only because the budget script carries it. Pre-load one quantified comparison per named competitor (vs. the catalog-size objection against Awardco specifically) and probe for the decision criteria instead of asserting the differentiation and moving on.
reporting-analytics 5 tests
q3-weighted-forecastPASS 1.00
Plain-English question: Build the Q3 2026 weighted forecast from the open pipeline, and show the math deal by deal.
Prompt — what the model saw
Below are the open deals in the default sales pipeline with a close date between 2026-07-01 and 2026-10-15, pulled from the CRM warehouse on 2026-09-05. Stages: DS1 (discovery) through DS5 (contract). Build the Q3 2026 forecast (quarter = 2026-07-01 to 2026-09-30). Weighting rule: 100 percent of COMMIT plus 35 percent of BEST_CASE. PIPELINE deals count zero. Only deals with a close date inside the quarter count. Report: COMMIT total, BEST_CASE total, weighted forecast, count of deals in each category inside the quarter, the deals you excluded for being outside the quarter (count and total amount), and the top 5 BEST_CASE deals by amount inside the quarter. Show arithmetic. Then, in a final '## Data quality' section of no more than four sentences, name the data quality problems in this extract that would make the forecast unreliable if it were run unattended.
cost $0.0055 · 80s · in 116 / out 6,144 tok
✓ commit_total_44729✓ best_case_total_203565✓ weighted_115976.75✓ commit_count_7✓ best_case_count_24✓ excluded_count_32✓ excluded_amount_227575✓ top5_aliases✓ flags_blank_owner✓ flags_missing_why_buys

✓ Correct Answer (ground truth)

Plain-English answer: A stage-weighted Q3 forecast number computed from the open pipeline, with per-stage math that reconciles to the deal data.
The answer in one line: excluded_amount = 227575; best_case_total = 203565; weighted_forecast = 115,976.75; commit_total = 44729; excluded_out_of_quarter = 32; best_case_count = 24; flags = blank owner column + empty why_buys fields must be surfaced (+2 more below)
excluded_amount227,575
best_case_count24
best_case_total203,565
commit_count7
commit_total44,729
excluded_out_of_quarter32
weighted_forecast115,976.75
flagsblank owner column + empty why_buys fields must be surfaced
top5_best_caseDeal-2D7423, Deal-25F752, Deal-E53952, Deal-5EED42, Deal-FA32A0
graded checkscommit_total_44729 · best_case_total_203565 · weighted_115976.75 · commit_count_7 · best_case_count_24 · excluded_count_32 · excluded_amount_227575 · top5_aliases · flags_blank_owner · flags_missing_why_buys

Model's actual answer

show the model's full answer
Q3 2026 FORECAST — default pipeline, quarter 2026-07-01 to 2026-09-30
Source: deals.csv extract pulled 2026-09-05. Weighting: 100% COMMIT + 35% BEST_CASE, PIPELINE = 0.

IN-QUARTER DEALS (close_date 2026-07-01 through 2026-09-30)

COMMIT — 7 deals
  Deal-547B2B  11,200  2026-09-11
  Deal-B7EBD1   9,000  2026-09-10
  Deal-403845   9,000  2026-09-11
  Deal-A2B47C   6,360  2026-09-11
  Deal-2465CE   5,400  2026-09-10
  Deal-A5E80A   2,520  2026-09-11
  Deal-499BF6   1,249  2026-09-30
  Sum: 11,200 + 9,000 + 9,000 + 6,360 + 5,400 + 2,520 + 1,249 = 44,729

BEST_CASE — 24 deals
  38,935 + 24,000 + 19,656 + 16,250 + 11,116 + 10,800 + 10,500 + 9,890 + 9,720 + 9,000 + 7,200 + 3,840 + 3,780 + 3,600 + 3,240 + 3,150 + 3,120 + 3,060 + 2,916 + 2,760 + 2,484 + 2,100 + 1,920 + 528
  = 203,565

PIPELINE — 23 deals, 0 in weighting (counts shown for completeness)

CATEGORY COUNTS INSIDE QUARTER
  COMMIT    7
  BEST_CASE 24
  PIPELINE  23

WEIGHTED FORECAST
  COMMIT:    44,729 x 1.00 = 44,729.00
  BEST_CASE: 203,565 x 0.35 = 71,247.75
    (203,565 x 0.30 = 61,069.50; 203,565 x 0.05 = 10,178.25; sum = 71,247.75)
  PIPELINE:  0
  TOTAL WEIGHTED FORECAST = 44,729.00 + 71,247.75 = 115,976.75

EXCLUDED — close date outside the quarter (2026-10-01 through 2026-10-15)
  Count: 32 deals
  Total amount: 227,575
  Sum: 43,875 + 18,000 + 17,000 + 13,770 + 10,800 + 9,000 + 9,000 + 7,920 + 7,690 + 7,500 + 7,200 + 5,700 + 5,400 + 5,400 + 5,400 + 5,400 + 5,160 + 4,800 + 4,400 + 4,300 + 4,000 + 3,600 + 3,600 + 3,600 + 3,300 + 2,400 + 1,800 + 1,800 + 1,680 + 1,600 + 1,400 + 1,080 = 227,575
  Includes 1 COMMIT (Deal-D348E1, 13,770, 2026-10-15) and 9 BEST_CASE. None fall before 2026-07-01, so no early-quarter exclusions.

TOP 5 BEST_CASE BY AMOUNT, INSIDE QUARTER
  1. Deal-2D7423   38,935   2026-09-30   DS3
  2. Deal-25F752   24,000   2026-09-25   DS4
  3. Deal-E53952   19,656   2026-09-30   DS4
  4. Deal-5EED42   16,250   2026-09-30   DS3
  5. Deal-FA32A0   11,116   2026-09-25   DS3
  (next highest in-quarter BEST_CASE is Deal-FC22A3 at 10,800)

## Data quality

The owner field is blank on 85 of 86 deals (only Deal-C9C286 carries an owner), so no deal can be attributed, chased, or held accountable. Forecast categories conflict with stages — Deal-A5E80A is COMMIT at DS1 and Deal-499BF6 is COMMIT at DS2, which inflates the 100%-weighted commit line with deals that have not reached contract. The extract contains past-dated open deals (Deal-333EBB closes 2026-08-28, Deal-57FF13 2026-09-02, Deal-7A2454 2026-09-04, all before the 2026-09-05 pull), indicating stale close dates not being maintained. Nearly all rows carry why_buys_chars = 0 (including every COMMIT deal), so the qualification basis behind the numbers is absent and the amounts cannot be independently validated.
aha-moment-2x20.83
Plain-English question: Do companies that hit 5+ givers and one redemption in month one actually retain better at 24 months?
Prompt — what the model saw
Test the activation hypothesis on the company cohort below: companies with 5+ unique givers (m1_users >= 5) AND 1+ successful redemption (m1_redemptions >= 1) in their first calendar month retain better at 24 months than those with one signal or neither. Every company in this extract is already 25+ months old, so current_status = 'active' means retained at 24 months. Report the full 2x2 (cohort size and 24-month retention rate for: both signals, givers-only, redemption-only, neither), state how many companies were excluded from the denominator and why (if none, say so), name the single signal with the largest retention lift, and state what this does and does not prove.
cost $0.0177 · 174s · in 14,489 / out 12,863 tok
✓ both_66✓ givers_47✓ redemp_31✓ neither_40✓ excluded_0✗ largest_lift_givers

✓ Correct Answer (ground truth)

Plain-English answer: The four cohort cells are graded against computed 24-month retention; the numbers come from the scorer's cohort table, not opinion.
The answer in one line: both_5plus_givers_and_redemption = 66.0% (31 of 47)
both_5plus_givers_and_redemption66.0% (31 of 47)
exclusionsnone — claim bounded as correlation, not causation
givers_only46.9% (23 of 49)
largest_single_signal_liftunique givers
neither40.0% (38 of 95)
redemption_only31.0% (9 of 29)
graded checksboth_66 · givers_47 · redemp_31 · neither_40 · excluded_0 · largest_lift_givers

Model's actual answer

show the model's full answer
ACTIVATION HYPOTHESIS TEST — 24-MONTH RETENTION (plg_company_cohort_2023.csv)

Universe: 220 companies, all signup_month 2023-01 through 2023-07 (every row is 25+ months old as of 2026-09). Retained = current_status == 'active'. Signal definitions as given: givers = m1_users >= 5; redemption = m1_redemptions >= 1.

EXCLUSIONS FROM THE DENOMINATOR: none. 0 companies excluded. Every row has non-null m1_users, m1_redemptions, and current_status; no company is younger than 25 months. All 220 rows are in the denominator. (Row count check: 47 + 49 + 29 + 95 = 220.)

FULL 2x2

                     n     retained   24-mo rate
both signals         47       31        31/47  = 66.0%
givers-only          49       23        23/49  = 46.9%
redemption-only      29        9         9/29  = 31.0%
neither              95       38        38/95  = 40.0%
TOTAL               220      101       101/220 = 45.9%

Lift vs "neither" (40.0% baseline):
- both:        66.0 - 40.0 = +26.0 pp
- givers-only: 46.9 - 40.0 =  +6.9 pp
- redemption-only: 31.0 - 40.0 = -9.0 pp (BELOW baseline)

SINGLE SIGNAL WITH THE LARGEST RETENTION LIFT

Marginal (each signal on/off, ignoring the other):

  m1_users >= 5:      54/96  = 56.25% retained
  m1_users < 5:       47/124 = 37.90% retained
  lift = 56.25 - 37.90 = +18.35 pp  (54+47 = 101 retained total; 96+124 = 220 — checks out)

  m1_redemptions >= 1: 40/76  = 52.63% retained
  m1_redemptions < 1:  61/144 = 42.36% retained
  lift = 52.63 - 42.36 = +10.27 pp  (40+61 = 101; 76+144 = 220 — checks out)

Answer: m1_users >= 5 (the 5+ unique givers signal), with a +18.35 pp marginal lift vs +10.27 pp for redemption.

WHAT THIS PROVES / DOES NOT PROVE

Proves (descriptively, in this extract):
- The activation hypothesis holds directionally for the "both" cell: 66.0% vs 40.0% for neither (+26.0 pp), and the ordering both > givers-only is consistent with givers being the stronger signal.
- Givers is the stronger of the two signals: +18.35 pp marginal vs +10.27 pp, and givers-only (46.9%) beats redemption-only (31.0%) by 15.9 pp.

Does NOT prove:
- Causation. This is a single retrospective cohort; no randomization, no control for company size, tier, industry, country, or concurrent onboarding effort. Companies that get 5+ givers in month 1 may simply be larger/more committed to start (selection). current_arr_usd and product_tier are heavily correlated with cell membership (all 'cancelled' rows are tier_two with ARR 0), so tier/ARR confounding is visible in the data itself.
- Redemption is not a standalone retention driver here. The redemption-only cell (31.0%) retains WORSE than "neither" (40.0%), so 1+ redemption without 5+ givers is not evidence of activation — it may mark small accounts with one engaged individual.
- No statistical significance testing was run (small cells: n=29 and n=47); the gaps could be noise at these sizes.
- 24-month retention only; no statement about ARR quality of retention (though note the largest ARR rows appear in the "neither" cell, e.g. C-0FF2DB67 at $6,156 and C-0CB2C1B4 at $7,896).
- Nothing about months beyond 24 or about causally driving retention by pushing givers/redemptions — the data cannot separate "activation causes retention" from "healthy companies activate."
arr-rollforward-reconciliationPASS 1.00
Plain-English question: Does CRM ARR match billing ARR — and if not, exactly where is the variance?
Prompt — what the model saw
Reconcile ARR between CRM company records and billing subscriptions as of 2026-09-05. Billing ARR is MRR x 12. Report both totals, the variance, and a decomposition into buckets (status mismatch, rounding, missing records, other) that sums exactly to the variance. List each mismatched account with a suggested owner. Business rule: any subscription with a term other than 12 months must have cf_agreement_end_date populated; list violations.
cost $0.0094 · 88s · in 9,646 / out 5,845 tok
✓ billing_total✓ crm_total✓ variance✓ buckets_sum✓ violations_listed

✓ Correct Answer (ground truth)

Plain-English answer: Billing $604,739.28 vs CRM $603,581.76: a $1,157.52 variance decomposed into status-mismatch, rounding, and missing-record buckets.
The answer in one line: n_companies = 40; n_mismatched_accounts = 7; as_of = 2026-09-05 (+4 more below)
n_companies40
n_mismatched_accounts7
as_of2026-09-05
billing_arr_total604739.28
bucket_sum_equals_variance1
buckets{"status_mismatch": {"amount": "-13158.48", "accounts": [{"company_alias": "C-0C8323BF", "billing_arr": "0.00", "hubspot_arr": "4905.24", "diff": "-4905.24", "suggested_owner": "Dana Mercer"}, {"compa
crm_arr_total603581.76
ruleterm_months != 12 requires cf_agreement_end_date
variance1157.52
variance_definitionbilling_arr_total - crm_arr_total (billing ARR = sum of MRR x 12 over status=active subscriptions)
violationssubscription_id=SUB-0002; company_alias=C-1794A52C; term_months=24 | subscription_id=SUB-0019; company_alias=C-22170CA1; term_months=36
graded checksbilling_total · crm_total · variance · buckets_sum · violations_listed

Model's actual answer

show the model's full answer
RECONCILIATION — CRM vs Billing ARR, as of 2026-09-05

CONVENTION (stated explicitly, since the data doesn't say):
- Billing ARR = MRR x 12 per subscription (per your rule).
- Cancelled subscriptions (SUB-000E, SUB-000F) generate no billing ARR, so the primary total EXCLUDES them. An all-subscription variant is shown at the end.
- Variance = CRM total − Billing total.

--------------------------------------------------
1. TOTALS (primary: active subscriptions only)
--------------------------------------------------
Billing ARR (37 active subs) = 617,897.76 − 4,905.24 (SUB-000E) − 8,253.24 (SUB-000F)
                             = 604,739.28
CRM ARR (39 company records) = 603,581.76
VARIANCE = 603,581.76 − 604,739.28 = −1,157.52

--------------------------------------------------
2. VARIANCE DECOMPOSITION (sums exactly to −1,157.52)
--------------------------------------------------
A. Missing records                     −11,952.00
   C-21629AA4: billing SUB-0004 = 2,370.77 x 12 = 28,449.24; no CRM record → −28,449.24
   C-0D5BBE3A: CRM 16,497.24; no billing subscription → +16,497.24
   Net: −28,449.24 + 16,497.24 = −11,952.00

B. Status mismatch                     +13,158.48
   SUB-000E (C-0C8323BF) cancelled: CRM still carries 408.77 x 12 = 4,905.24 → +4,905.24
   SUB-000F (C-0DC4FB8C) cancelled: CRM still carries 687.77 x 12 = 8,253.24 → +8,253.24
   (CRM includes ARR that cancelled billing subs no longer bill.)

C. Rounding                                 +36.00
   C-0D66DF9E: 1,932.00 x 12 = 23,184.00 vs CRM 23,200.00 → +16.00
   C-14D70CE0: 1,515.00 x 12 = 18,180.00 vs CRM 18,200.00 → +20.00
   (CRM values are round hundreds — rounded/entered values, not cent-level artifacts. Classified as rounding by judgment; if you treat them as unexplained, they move to "other.")

D. Other                                −2,400.00
   C-0F7269D7: 2,233.00 x 12 = 26,796.00 vs CRM 24,396.00 → −2,400.00 (real delta, cause unknown from data)

SUM: −11,952.00 + 13,158.48 + 36.00 − 2,400.00 = −1,157.52 ✓

--------------------------------------------------
3. MISMATCHED ACCOUNTS + SUGGESTED OWNER
--------------------------------------------------
C-21629AA4  billing 28,449.24 / CRM none      −28,449.24  RevOps — CRM data steward: create the company record (SUB-0004 active)
C-0D5BBE3A  billing none / CRM 16,497.24      +16,497.24  Billing ops — locate missing subscription or RevOps zeroes CRM ARR
C-0C8323BF  SUB-000E cancelled / CRM 4,905.24  +4,905.24  AM/CSM — confirm churn; RevOps zero the CRM ARR
C-0DC4FB8C  SUB-000F cancelled / CRM 8,253.24  +8,253.24  AM/CSM — confirm churn; RevOps zero the CRM ARR
C-0F7269D7  billing 26,796.00 / CRM 24,396.00  −2,400.00  Deal desk — reconcile against order form (possible mid-term change not synced)
C-0D66DF9E  billing 23,184.00 / CRM 23,200.00     +16.00  RevOps — reset CRM ARR to 23,184.00 / adopt unrounded convention
C-14D70CE0  billing 18,180.00 / CRM 18,200.00     +20.00  RevOps — reset CRM ARR to 18,180.00 / adopt unrounded convention

All 32 other accounts match exactly (diff 0.00).

--------------------------------------------------
4. TERM-RULE VIOLATIONS (term ≠ 12 months, cf_agreement_end_date empty)
--------------------------------------------------
SUB-0002  C-1794A52C  24 months  active  VIOLATION (field blank)
SUB-0019  C-22170CA1  36 months  active  VIOLATION (field blank)

Compliant non-12-month subs (for contrast): SUB-000C (C-0DB48281, 24mo, 2027-11-30), SUB-001A (C-0FC4DBB8, 36mo, 2027-11-30).

--------------------------------------------------
5. VARIANT — billing ARR counting cancelled subs too (literal MRR x 12 on all 39)
--------------------------------------------------
Billing total = 617,897.76; CRM = 603,581.76; variance = −14,316.00
Buckets: missing −11,952.00; status mismatch 0.00 (cancelled matched both sides); rounding +36.00; other −2,400.00. Sum = −14,316.00 ✓

Data notes: no field links subs to CRM ARR source-of-truth conventions; "suggested owner" assignments are workflow judgment, not from the data. Cause of the C-0F7269D7 −2,400.00 delta cannot be determined from the two files provided.
kvm-month-over-monthPASS 1.00
Plain-English question: How did the four core KVMs move in August versus July?
Prompt — what the model saw
Report the four core KVMs for 2026-08 against 2026-07: giving rate, redemptions per user, 1:1 meetings engagement, pulse check engagement. For each: value, prior value, absolute and relative change, direction. Name the KVM with the largest relative move and the segment (plan_tier or size_band) driving it if the data supports that; otherwise say the data does not. Under 200 words plus one table.
cost $0.0081 · 48s · in 10,171 / out 4,120 tok
✓ four_kvms✓ biggest_mover_named✓ direction_correct✓ segment_named✓ word_cap

✓ Correct Answer (ground truth)

Plain-English answer: Giving rate 60.27% (up 0.07%), redemptions per user 1.7302 (up 0.01%), with 1:1 meetings and pulse engagement each up fractionally — direction reported for all four.
The answer in one line: n_company_rows_per_month = 30; aggregation = simple mean of company-level values per month
n_company_rows_per_month30
aggregationsimple mean of company-level values per month
biggest_moverpulse_engagement
driving_segmentsize_band=enterprise; 2026-07=0.55; 2026-08=0.2743; rel_change_pct=-50.13
kvms{"giving_rate": {"2026-07": 0.6023, "2026-08": 0.6027, "abs_change": 0.0004, "rel_change_pct": 0.07, "direction": "up"}, "redemptions_per_user": {"2026-07": 1.73, "2026-08": 1.7302, "abs_change": 0.00
segment_detail{"smb": {"2026-07": 0.6588, "2026-08": 0.6573, "rel_change_pct": -0.23}, "mid_market": {"2026-07": 0.593, "2026-08": 0.5942, "rel_change_pct": 0.2}, "enterprise": {"2026-07": 0.55, "2026-08": 0.2743,
graded checksfour_kvms · biggest_mover_named · direction_correct · segment_named · word_cap

Model's actual answer

show the model's full answer
KVM means are unweighted averages across 30 companies (10 per size_band; all plan_tier = tier_three).

| KVM | 2026-07 | 2026-08 | Abs Δ | Rel Δ | Direction |
|---|---|---|---|---|---|
| Giving rate | 0.6023 | 0.6027 | +0.0004 | +0.07% | up (flat) |
| Redemptions/user | 1.7300 | 1.7302 | +0.0002 | +0.01% | up (flat) |
| 1:1 meetings engagement | 0.4469 | 0.4472 | +0.0003 | +0.07% | up (flat) |
| Pulse check engagement | 0.6006 | 0.5086 | −0.0920 | −15.31% | down |

Arithmetic (example): pulse 08 mean = sum of 30 pulse_engagement values ÷ 30 = 15.2583 ÷ 30 = 0.508610; 07 = 18.0176 ÷ 30 = 0.600587; Δ = −0.091977; rel = −0.091977 / 0.600587 = −15.31%. Others computed identically.

Largest relative move: pulse check engagement, −15.3%.

Driving segment: size_band = enterprise. Enterprise pulse fell 0.54998 → 0.27428 (−0.2757, −50.1%, n=10), while smb went 0.65879 → 0.65731 (−0.0015) and mid_market 0.59299 → 0.59424 (+0.0013). The enterprise drop accounts for the entire portfolio-level decline; every enterprise alias (C-0B2895EF, C-0B2213A9, C-0D6CC8E3, C-0D0B047C, C-0D3278C7, C-0FCCD2DF, C-0F6C0F34, C-8C2E8F00, C-0B827671, C-0BA71F12) roughly halved. plan_tier cannot discriminate (all tier_three).
rewards-ytd-sectionPASS 1.00
Plain-English question: Produce the redemption section of the weekly rewards report through August.
Prompt — what the model saw
Produce the redemption section of the weekly rewards report through the last completed month (2026-08): redemption count, spend, unique redeemers, redemptions per redeemer, provider mix as percent of spend, top 5 countries by redemptions. Say which month is the last completed month. Provider shares must sum to 100.
cost $0.0416 · 239s · in 47,345 / out 23,707 tok
✓ count✓ spend✓ unique_redeemers✓ month_named✓ provider_mix✓ top_country

✓ Correct Answer (ground truth)

Plain-English answer: 378 redemptions, $27,846 spend, 235 unique redeemers at 1.61 each, custom rewards at 39% of spend, US leading all countries.
The answer in one line: redemption_count = 378; unique_redeemers = 235; redemptions_per_redeemer = 1.61; total_spend_usd = 27846.00
redemption_count378
redemptions_per_redeemer1.61
unique_redeemers235
provider_mix_pct_of_spendTangoCard=11.6; Tremendous=30.54; custom=39.05; Snappy=18.81
top5_countries_by_countcountry=US; count=244 | country=CA; count=24 | country=AU; count=21 | country=GB; count=17 | country=NL; count=17
total_spend_usd27846.00
graded checkscount · spend · unique_redeemers · month_named · provider_mix · top_country

Model's actual answer

show the model's full answer
REDEMPTION SECTION — WEEKLY REWARDS REPORT (YTD through last completed month)

Last completed month: 2026-08 (data max = 2026-08-31T11:53:00; today is 2026-09-24, so 2026-09 is incomplete and excluded — no 2026-09 rows exist in the file anyway). All 378 rows fall between 2026-01-01 and 2026-08-31, so the full file is in scope and nothing is dropped. Source: redemptions_ytd.csv (378 data rows).

1) Redemption count
   378 redemptions (2026-01: 45, 02: 36, 03: 50, 04: 52, 05: 43, 06: 46, 07: 58, 08: 48; sum = 378)

2) Spend
   $27,846.00 USD
   Cross-check by provider: 10,873 (custom) + 8,505 (Tremendous) + 5,238 (Snappy) + 3,230 (TangoCard) = 27,846

3) Unique redeemers
   235 distinct user_key values

4) Redemptions per redeemer
   378 / 235 = 1.6085 → 1.61 redemptions per redeemer

5) Provider mix (% of spend)
   custom       $10,873.00 / $27,846.00 = 39.05%   (37 redemptions)
   Tremendous    $8,505.00 / $27,846.00 = 30.54%   (192 redemptions)
   Snappy        $5,238.00 / $27,846.00 = 18.81%   (59 redemptions)
   TangoCard     $3,230.00 / $27,846.00 = 11.60%   (90 redemptions)
   Sum: 39.05 + 30.54 + 18.81 + 11.60 = 100.00%
   (unrounded shares: 39.0469 + 30.5430 + 18.8106 + 11.5995 = 100.0000)

6) Top 5 countries by redemptions (count)
   US: 244
   CA: 24
   AU: 21
   GB: 17
   NL: 17
   Note: GB and NL tie at 17 redemptions — both are shown as 5th; the next country is SG at 12. Country counts sum: 244+24+21+17+17 = 323 of 378 (remaining: SG 12, DE 9, FR 9, CH 9, IN 8, BR 6, NZ 2 = 55; 323 + 55 = 378).

Data gaps / caveats: none affecting the above — the file covers 2026-01 through 2026-08 only, so this is a YTD-through-August view, not a rolling-12-month view. No data is missing for the requested metrics.
customer-success 5 tests
churn-save-eligibilityPASS 1.00
Plain-English question: Which at-risk accounts actually qualify for a churn-save offer, and which play fits each?
Prompt — what the model saw
Which at-risk accounts qualify for a churn-save offer under the documented eligibility rules, what amount is at stake per account and in total, and which play fits each (usage revival, executive touch, commercial concession)? Cite the signal that justifies each play. List accounts that look at risk but do not qualify and why.
cost $0.0071 · 124s · in 4,736 / out 5,680 tok
✓ eligible_set✓ total_at_stake✓ plays_cited✓ noneligible_named✓ rules_applied

✓ Correct Answer (ground truth)

Plain-English answer: Only accounts meeting all three rules qualify — e.g. C-0F6C0F34 at $49,707 (executive touch) and C-0B827671 at $25,365 (usage revival) — with amount and play per account.
The answer in one line: n_accounts = 30; snapshot_date = 2026-09-05
n_accounts30
eligibleaccount_alias=C-0F6C0F34; amount_at_stake=49707.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0B827671; amount_at_stake=25365.00; play=usage revival; justifying_signal=usage_trend_3m=declining over the last 3 months | account_alias=C-0B360C78; amount_at_stake=35748.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 75% - risk is commercial, not adoption | account_alias=C-0B0F1BAB; amount_at_stake=5494.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0CA21961; amount_at_stake=16829.00; play=usage revival; justifying_signal=seat utilization 26% is below 50% | account_alias=C-0E9C27D1; amount_at_stake=41235.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 85% - risk is commercial, not adoption … (+2 more)
non_eligible_at_riskaccount_alias=C-0BC71BDD; health_score=55 | account_alias=C-0BA71F12; health_score=52 | account_alias=C-0F6694C3; health_score=43 | account_alias=C-0BE96399; health_score=54 | account_alias=C-0F876796; health_score=47 | account_alias=C-0FCCD2DF; health_score=43 … (+1 more)
ruleshealth_score < 60, churn_save_eligible_amount > 0, renewal within 120 days of snapshot
snapshot_date2026-09-05
total_amount_at_stake224601.00
graded checkseligible_set · total_at_stake · plays_cited · noneligible_named · rules_applied

Model's actual answer

show the model's full answer
ELIGIBILITY SCREEN (rules R1+R2+R3, all must pass)
Snapshot date: 2026-09-05. R3 window: 2026-09-05 + 120 days = 2027-01-03 (renewal_date must be on/before that).
R1: health_score < 60. R2: churn_save_eligible_amount > 0. R3: renewal within 120 days.

QUALIFYING ACCOUNTS (8)

Alias        Health  Eligible $   ARR       Renewal      Days out  Play                Justifying signal
C-0F6C0F34     51    49,707.00   86,741.00  2026-10-03      28     Executive touch     champion_active=false (usage growing, 308/395 seats = 78%)
C-0B827671     56    25,365.00   72,088.00  2026-11-14      70     Usage revival       usage_trend_3m=declining (113/202 seats = 56%)
C-0B360C78     57    35,748.00   60,427.00  2026-10-28      53     Commercial conces.  no adoption/relationship signal: usage growing (246/327 = 75%), champion_active=true — risk is unexplained by usage/relationship
C-0B0F1BAB     38     5,494.00   15,391.00  2026-09-23      18     Executive touch     champion_active=false (also lowest health at 38; usage flat)
C-0CA21961     58    16,829.00   31,501.00  2026-12-28     114     Usage revival       seats_used 84 of 325 = 25.8% utilization (flat trend)
C-0E9C27D1     39    41,235.00   75,093.00  2026-09-24      19     Commercial conces.  usage flat at high utilization (134/157 = 85%), champion_active=true — no adoption/relationship gap; renewal in 19 days
C-0CEF69FD     53    32,621.00   79,324.00  2026-11-21      77     Executive touch     champion_active=false (usage growing, 97/136 = 71%)
C-0D3278C7     54    17,602.00   33,815.00  2026-11-12      68     Usage revival       usage_trend_3m=declining (126/380 seats = 33.2%)

AMOUNT AT STAKE
Per account (churn_save_eligible_amount):
49,707.00 + 25,365.00 + 35,748.00 + 5,494.00 + 16,829.00 + 41,235.00 + 32,621.00 + 17,602.00
Step: 49,707 + 25,365 = 75,072
      75,072 + 35,748 = 110,820
      110,820 + 5,494 = 116,314
      116,314 + 16,829 = 133,143
      133,143 + 41,235 = 174,378
      174,378 + 32,621 = 206,999
      206,999 + 17,602 = 224,601
TOTAL AT STAKE: $224,601.00
(For reference, combined ARR of these 8: 434,680.00 — 86,741+72,088+60,427+15,391+31,501+75,093+79,324+33,815 = 434,680.)

Play totals: Usage revival 59,856.00 (25,365+16,829+17,602); Executive touch 87,822.00 (49,707+5,494+32,621); Commercial concession 76,983.00 (35,748+41,235). Check: 59,856+87,822+76,983 = 224,601. ✓

NOTE ON PLAY MAPPING: eligibility_rules.csv defines only eligibility (R1–R3), not a play policy. The play assignments above are my mapping from the signals present in churnzero_accounts.csv, using this stated precedence: (1) usage revival if usage_trend_3m=declining or seat utilization < 40%; (2) else executive touch if champion_active=false; (3) else commercial concession (no adoption or relationship signal present). If a documented play policy exists elsewhere, it was not provided and should override this.

AT-RISK BUT NOT QUALIFYING (health < 60, failed at least one rule)

Alias        Health  Eligible $   Renewal      Failed rule(s)                          Why
C-0BC71BDD     55        0.00    2026-10-27    R2 (amount = 0)                          At risk and renewal in window (52 days), but churn_save_eligible_amount is 0.00 — nothing at stake per R2
C-0BE96399     54        0.00    2026-10-29    R2 (amount = 0)                          Declining usage and renewal in 54 days, but eligible amount is 0.00 — closest miss alongside C-0BC71BDD
C-10A56B0F     54        0.00    2026-12-12    R2 (amount = 0)                          Declining usage, but eligible amount is 0.00
C-0F876796     47   19,958.00    2027-02-06    R3 (154 days > 120)                      Has $19,958 eligible and declining usage, but renewal is 154 days out — outside the 120-day window (fails by 34 days)
C-0BA71F12     52    6,824.00    2027-04-11    R3 (218 days > 120)                      Renewal 218 days out; ineligible on timing despite $6,824 eligible and declining usage (23/98 seats = 23%)
C-0F6694C3     43        0.00    2027-03-21    R2 (amount = 0) and R3 (197 days)          Fails both: no eligible amount and renewal 197 days out
C-0FCCD2DF     43        0.00    2027-04-23    R2 (amount = 0) and R3 (230 days)          Fails both: no eligible amount and renewal 230 days out

All remaining accounts have health_score ≥ 60 and fail R1 (not at risk). Arithmetic note: no row with health < 60 and amount > 0 has a renewal date between 2027-01-04 and 2027-02-05, so the two failing sets (R2-only and R3-only) do not overlap except where both flags are listed above.
expansion-kitPASS 1.00
Plain-English question: Build the expansion kit for the customer who just replied to our outreach.
Prompt — what the model saw
A customer replied to outreach (reply included). Build the expansion kit: seat coverage (licensed over headcount), two lines on usage health, headroom in seats and ARR at the current per-seat rate, who replied and whether they can buy, the right buyer from our contacts, and a reply email under 150 words that moves toward the right person without being pushy and cites one real usage fact.
cost $0.0025 · 26s · in 3,454 / out 1,053 tok
✓ coverage_pct✓ headroom_arr✓ right_buyer✓ replier_not_buyer✓ usage_fact✓ email_under_150

✓ Correct Answer (ground truth)

Plain-English answer: 37.5% seat coverage, 250 headroom seats worth $15,000 ARR at $60 a seat, usage up 43.2% in six months — and the replier cannot buy; Dana R. (VP People) is the right buyer.
The answer in one line: headroom_seats = 250; coverage_pct = 37.5; company_alias = C-0DDFC9A7 (+3 more below)
headroom_seats250
coverage_pct37.5
company_aliasC-0DDFC9A7
headroom_arr15000.00
per_seat_annual_rate60.00
replierMaria S.
replier_can_buy0
right_buyername=Dana R.; title=VP People
usage_factsactive_users_2026_08=126; active_users_2026_03=88; growth_pct_6m=43.2; licensed_seat_utilization_pct=84
usage_health_linesMonthly active users grew from 88 (2026-03) to 126 (2026-08), +43.2% over 6 months., Last month 126 of 150 licensed seats were active (84.0% utilization), approaching the seat cap.
graded checkscoverage_pct · headroom_arr · right_buyer · replier_not_buyer · usage_fact · email_under_150

Model's actual answer

show the model's full answer
EXPANSION KIT — C-0DDFC9A7

SEAT COVERAGE
Licensed 150 / headcount 400 = 37.5% of the org covered.
Unlicensed headcount = 400 − 150 = 250 seats.

USAGE HEALTH
1. MAU grew every month, 6 straight months: 88 → 95 → 102 → 110 → 118 → 126. Change = 126 − 88 = +38 users = +43.2% (38/88).
2. Activation is high and rising: 126/150 = 84.0% of licensed seats active in Aug 2026, up from 88/150 = 58.7% in Mar 2026. Only 150 − 126 = 24 licensed seats unused.

HEADROOM (at current per-seat rate)
Per-seat rate = $9,000.00 / 150 = $60.00 per seat (ARR basis).
- Seat headroom: 250 unlicensed seats (to full headcount coverage). At current MAU trajectory there's no slack in licenses — only 24 seats spare.
- ARR headroom: 250 × $60.00 = $15,000.00 incremental ARR if licensed to full headcount. Full-coverage ARR would be 400 × $60.00 = $24,000.00 (vs $9,000.00 today = 2.7x).

WHO REPLIED / BUYING AUTHORITY
Maria S., People Operations Coordinator (replied 2026-09-02). Explicitly not a purchasing decision-maker — she states budget and seat expansion sit with Dana R. and offers an introduction. Not a buyer, but a willing internal champion (team "genuinely loves Bonusly").

RIGHT BUYER
Dana R., VP People — owns budget and seat expansion per Maria. Last engaged 2026-05-18 (~3.5 months stale). Note from the reply: Dana "has been asking about our usage numbers lately" — warm context for outreach. (Sam K., Office Manager, last engaged 2025-11-03, is not relevant to this purchase.)

REPLY EMAIL (128 words)

Subject: Re: Growing your team's recognition program — intro to Dana?

Hi Maria,

Thanks for the quick reply — glad the team's enjoying Bonusly.

That's helpful context, and no problem at all on the purchasing side. If you're comfortable making an introduction to Dana R., I'd appreciate it. One data point that may be useful for the conversation she's been having about usage: your monthly active users grew from 88 in March to 126 in August — six straight months of growth, now 84% of your 150 licensed seats.

Happy to share a simple usage summary Dana can use, or to walk her through it directly — whatever's easiest for you.

Thanks again for pointing me in the right direction.

Best,
Cole Ingram

DATA NOTE: All figures derived solely from expansion_account.csv and the reply. No expansion discounting, no assumed pricing tier change — ARR headroom assumes flat $60/seat.
onboarding-statusPASS 1.00
Plain-English question: What is done and what is missing in this customer's onboarding before the mid-onboarding call?
Prompt — what the model saw
Prep the CSM for the mid-onboarding call: what is complete (integrations connected, allowance set, admins added, first recognition given, first redemption), what is not, early engagement signals, and the three things to cover on the call. Mark nothing complete without a data field that shows it.
cost $0.0037 · 46s · in 3,541 / out 2,388 tok
✓ complete_items✓ incomplete_items✓ three_items✓ grounded

✓ Correct Answer (ground truth)

Plain-English answer: 25 days live: Slack connected, allowance set, admins added — but the HRIS integration is not connected, with early engagement signals and three call priorities.
The answer in one line: days_live = 25; company_alias = C-0D284E42
days_live25
call_itemsConnect the HRIS integration (integration_hris is empty) to automate user provisioning and offboarding., Get to first redemption (first_redemption_at is empty): walk admins through the rewards catalog and confirm allowance funding., Build the giving habit: daily active givers averaged 4.3 in week 1 vs 13.0 last week - agree on manager nudges and a recognition cadence.
checklist{"slack_integration_connected": {"complete": true, "field": "integration_slack", "value": "2026-08-12"}, "hris_integration_connected": {"complete": false, "field": "integration_hris", "value": ""}, "a
company_aliasC-0D284E42
engagement_signalsavg_active_givers_week1=4.3; avg_active_givers_last_week=13; latest_active_givers=15; days_with_activity=25
graded checkscomplete_items · incomplete_items · three_items · grounded

Model's actual answer

show the model's full answer
CSM PREP — C-0D284E42 mid-onboarding call
Source: onboarding_account.csv, onboarding_usage.csv (usage ends 2026-09-04; nothing known after that date)

WHAT IS COMPLETE (each backed by a data field)
- Integrations — Slack only: integration_slack = 2026-08-12 (+1 day after signup 2026-08-11)
- Allowance set: allowance_set = 2026-08-13 (+2 days after signup)
- Admins added: admins_added = 2
- First recognition: first_recognition_at = 2026-08-15 14:22 (+4 days after signup)

WHAT IS NOT COMPLETE (field empty — treat as not done)
- HRIS integration: integration_hris is blank. Not connected.
- First redemption: first_redemption_at is blank. As of the last data point (2026-09-04) that is 22 days with allowance set (2026-08-13 → 2026-09-04) and zero redemptions.

EARLY ENGAGEMENT SIGNALS (active_givers, 25 days: 2026-08-11 → 2026-09-04)
- Trajectory: 3 → 15 active givers. Delta +12, 5.0x (15 ÷ 3). Peak 15 on 2026-09-03 and 2026-09-04.
- Weekly averages (sum ÷ days):
  - W1 08-11→08-17: 3+3+4+4+5+4+7 = 30 / 7 = 4.3
  - W2 08-18→08-24: 5+7+6+9+8+9+9 = 53 / 7 = 7.6
  - W3 08-25→08-31: 9+11+10+10+11+13+11 = 75 / 7 = 10.7
  - W4 09-01→09-04: 13+13+15+15 = 56 / 4 = 14.0 (partial week)
- Every day has ≥3 active givers (min 3 on 08-11 and 08-12) — no dead days in the window.
- No dip below the prior week's average at any week boundary — consistent upward trend, not a launch spike.

DATA GAPS (say so on the call; do not guess)
- No total-employee/user count provided → adoption % cannot be computed. "15 active givers" is unbounded without headcount.
- active_givers is a per-day figure; cumulative unique givers cannot be derived (values can't be summed without double-counting).
- No usage data after 2026-09-04 — current engagement (as of the call date) is unknown.

THREE THINGS TO COVER ON THE CALL
1. HRIS integration — field is blank. Confirm whether it's blocked, in flight, or deprioritized; agree owner and target date.
2. First redemption — allowance has been live since 2026-08-13 with zero redemptions recorded. Diagnose why (e.g., is the reward catalog visible to givers/recipients?) and agree a "first redemption by" milestone.
3. Sustain + scope the growth — daily givers grew 5x to 15, but without headcount we can't size penetration. Get employee count, agree an adoption target, and confirm current engagement since 2026-09-04 (data ends there).
renewal-risk-conflicting-datesPASS 1.00
Plain-English question: Two systems disagree on renewal dates — which do you trust per account, and who is actually renewing inside 90 days?
Prompt — what the model saw
Produce a 90-day renewal risk brief. Two systems hold renewal dates and disagree on some accounts; decide which to trust per account and say why (multi-year contracts are known to be wrong in ChurnZero). For every renewal: company, CSM, ARR, date used, seat utilization, 3-month usage trend, risk rating with one sentence of evidence. Flag every disagreement. Close with total ARR renewing and ARR at risk.
cost $0.0373 · 352s · in 30,930 / out 27,059 tok
✓ total_renewing✓ arr_at_risk✓ disagreements_flagged✓ trust_rule

✓ Correct Answer (ground truth)

Plain-English answer: Chargebee wins per account (multi-year contracts are known wrong in ChurnZero), with each trusted date, the reason, and a 90-day-window flag.
The answer in one line: n_accounts = 20; n_disagreements = 5; snapshot_date = 2026-09-05 (+2 more below)
n_accounts20
n_disagreements5
accounts20 items — e.g. account_alias=C-0B144C78; csm=Cole Ingram; arr=30899.00; trusted_renewal_date=2026-11-02; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.4; usage_3m_ratio=1.03; risk=low; evidence=3-month usage ratio 1.03 (last3 avg 103 vs prior3 100), seat utilization 75% | account_alias=C-0B20DB64; csm=Dana Mercer; arr=21770.00; trusted_renewal_date=2026-10-07; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=56.6; usage_3m_ratio=1; risk=medium; evidence=3-month usage ratio 1.00 (last3 avg 295 vs prior3 295), seat utilization 57% | account_alias=C-0B344485; csm=Elena Sinclair; arr=64384.00; trusted_renewal_date=2026-11-16; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=78; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 241 vs prior3 231), seat utilization 78% | account_alias=C-0B369871; csm=Dana Mercer; arr=85128.00; trusted_renewal_date=2026-10-29; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.1; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 330 vs prior3 316), seat utilization 75%
arr_at_risk_90d359409.00
disagreementsaccount_alias=C-0B7D2C30; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-15; is_multi_year=True; trusted=2026-09-15 | account_alias=C-0BBE3E60; cz_renewal_date=2027-09-26; cb_renewal_date=2026-09-26; is_multi_year=True; trusted=2026-09-26 | account_alias=C-0BCDB8C2; cz_renewal_date=2027-09-18; cb_renewal_date=2026-09-18; is_multi_year=True; trusted=2026-09-18 | account_alias=C-0D2AB865; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-22; is_multi_year=True; trusted=2026-09-22 | account_alias=C-0F5D2323; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-29; is_multi_year=True; trusted=2026-09-29
snapshot_date2026-09-05
total_arr_renewing_90d1048715.00
trust_rulemulti-year contracts: Chargebee is authoritative (ChurnZero known wrong); otherwise systems agree or Chargebee wins
window2026-09-05 to 2026-12-04
graded checkstotal_renewing · arr_at_risk · disagreements_flagged · trust_rule

Model's actual answer

show the model's full answer
90-DAY RENEWAL RISK BRIEF
As-of: 2026-09-24 (no as-of date supplied in the data; run date used = session date 2026-09-24)
Window: 2026-09-24 → 2026-12-23. Trusted renewal dates in the provided book span 2026-09-15 → 2026-11-24. All 20 accounts in the files are covered; 3 carry trusted dates 1–9 days BEFORE the as-of date and are flagged PAST DUE (elapsed date — the data contains no renewal outcome).

SOURCE-OF-TRUTH DECISION (per account)
Rule applied: multi-year contracts (chargebee_terms.csv is_multi_year=true) are known-wrong in ChurnZero → use Chargebee date. Single-year contracts where both systems agree → date is unambiguous (either source).
- 5 accounts disagree, and all 5 are multi-year → Chargebee trusted in every disagreement (see FLAG list at the end).
- 15 accounts agree exactly (12-month terms) → no decision needed.
Supporting pattern: ChurnZero shows 2026-09-10 for 3 accounts where Chargebee says 09-15/09-22/09-29 — consistent with a default/import artifact, not contract reality.

METHOD (arithmetic shown per row)
- Seat utilization = seats_used / seats.
- 3-month usage trend = (active_users 2026-08 − active_users 2026-06) / active_users 2026-06. Monthly values shown as Jun/Jul/Aug.
- Rubric: HIGH = 3-mo trend ≤ −10%, or seat utilization < 40% with flat/declining usage. MEDIUM = trend −10% to −3%, or utilization 40–60% with declining usage, or Aug actives < 40% of seats_used (provisioning far above real engagement = renewal-shrink exposure). LOW = trend > −3% and utilization ≥ 50% with no engagement gap.

RENEWALS (ordered by date used)

1. C-0B7D2C30 | CSM: Dana Mercer | ARR $65,901.00
   Date used: 2026-09-15 (Chargebee; CZ says 2026-09-10 — DISAGREEMENT, 36-mo multi-year → CZ untrusted). PAST DUE.
   Seat utilization: 274/476 = 57.6% | 3-mo usage: 97 / 94 / 84 = −13.4% (Jun→Aug)
   RISK: HIGH — Usage has fallen every month for 12 straight months (155 → 84, −45.8% YoY) and the trusted renewal date already passed.

2. C-0BCDB8C2 | CSM: Cole Ingram | ARR $54,427.00
   Date used: 2026-09-18 (Chargebee; CZ says 2027-09-18 — DISAGREEMENT, full year off, 36-mo multi-year → CZ untrusted). PAST DUE.
   Seat utilization: 232/424 = 54.7% | 3-mo usage: 127 / 118 / 110 = −13.4% (Jun→Aug)
   RISK: HIGH — Twelve consecutive months of decline (200 → 110, −45.0% YoY) with only 54.7% of seats used and the renewal date already elapsed.

3. C-0D2AB865 | CSM: Elena Sinclair | ARR $38,022.00
   Date used: 2026-09-22 (Chargebee; CZ says 2026-09-10 — DISAGREEMENT, 24-mo multi-year → CZ untrusted). PAST DUE.
   Seat utilization: 250/407 = 61.4% | 3-mo usage: 125 / 117 / 109 = −12.8% (Jun→Aug)
   RISK: HIGH — Usage is down 45.2% YoY (199 → 109) with the decline still running through August, and the trusted renewal date has passed.

4. C-0BBE3E60 | CSM: Dana Mercer | ARR $30,993.00
   Date used: 2026-09-26 (Chargebee; CZ says 2027-09-26 — DISAGREEMENT, full year off, 24-mo multi-year → CZ untrusted).
   Seat utilization: 74/114 = 64.9% | 3-mo usage: 39 / 35 / 33 = −15.4% (Jun→Aug)
   RISK: HIGH — Steepest 3-month drop in the book (−15.4%) on top of −47.6% YoY (63 → 33), with renewal 2 days out.

5. C-0F5D2323 | CSM: Cole Ingram | ARR $90,647.00
   Date used: 2026-09-29 (Chargebee; CZ says 2026-09-10 — DISAGREEMENT, 24-mo multi-year → CZ untrusted).
   Seat utilization: 111/390 = 28.5% | 3-mo usage: 20 / 21 / 18 = −10.0% (Jun→Aug)
   RISK: HIGH — Largest single exposure in the book runs on 28.5% seat utilization and 18 active users (4.6% of 390 seats, 16.2% of seats used) at $90,647 — clear downsell exposure at a renewal 5 days out.

6. C-0EC6999D | CSM: Elena Sinclair | ARR $79,419.00
   Date used: 2026-10-03 (both systems agree; 12-mo single-year)
   Seat utilization: 31/112 = 27.7% | 3-mo usage: 17 / 16 / 15 = −11.8% (Jun→Aug)
   RISK: HIGH — Only 31 of 112 seats used and ~15 monthly actives supporting $79,419, with usage flat in a 14–17 band all year (YoY 0.0%).

7. C-0B20DB64 | CSM: Dana Mercer | ARR $21,770.00
   Date used: 2026-10-07 (both systems agree; 12-mo single-year)
   Seat utilization: 214/378 = 56.6% | 3-mo usage: 294 / 298 / 294 = 0.0% (Jun→Aug)
   RISK: LOW — Usage is flat in a 293–298 band for 12 months (YoY +0.3%), a stable base ahead of renewal.

8. C-0BBC4E7A | CSM: Cole Ingram | ARR $56,374.00
   Date used: 2026-10-10 (both systems agree; 12-mo single-year)
   Seat utilization: 228/337 = 67.7% | 3-mo usage: 142 / 141 / 139 = −2.1% (Jun→Aug)
   RISK: LOW — Usage sits in a 139–142 band all year (−2.1% over 3 months) with 67.7% of seats used.

9. C-0FD551AB | CSM: Elena Sinclair | ARR $48,815.00
   Date used: 2026-10-14 (both systems agree; 12-mo single-year)
   Seat utilization: 210/376 = 55.9% | 3-mo usage: 123 / 122 / 126 = +2.4% (Jun→Aug)
   RISK: LOW — Usage is flat-to-up across 12 months (124 → 126, YoY +1.6%) with 55.9% seat utilization.

10. C-0F9F8F13 | CSM: Dana Mercer | ARR $46,230.00
    Date used: 2026-10-18 (both systems agree; 12-mo single-year)
    Seat utilization: 199/352 = 56.5% | 3-mo usage: 185 / 185 / 182 = −1.6% (Jun→Aug)
    RISK: LOW — Usage has held a 181–185 band for 12 months (YoY 0.0%) at 56.5% seat utilization.

11. C-0BC34584 | CSM: Cole Ingram | ARR $16,740.00
    Date used: 2026-10-22 (both systems agree; 12-mo single-year)
    Seat utilization: 327/494 = 66.2% | 3-mo usage: 104 / 104 / 106 = +1.9% (Jun→Aug)
    RISK: MEDIUM — Usage is slightly up (YoY +2.9%) but only 106 of 327 seats-used were active in August (32.4%), leaving shrink exposure at renewal.

12. C-0B7A7546 | CSM: Elena Sinclair | ARR $35,062.00
    Date used: 2026-10-25 (both systems agree; 12-mo single-year)
    Seat utilization: 182/205 = 88.8% | 3-mo usage: 64 / 65 / 63 = −1.6% (Jun→Aug)
    RISK: MEDIUM — 88.8% of seats are marked used yet just 63 users were active in August (34.6% of seats-used) despite +8.6% YoY growth, so provisioning runs far ahead of real engagement.

13. C-0B369871 | CSM: Dana Mercer | ARR $85,128.00
    Date used: 2026-10-29 (both systems agree; 12-mo single-year)
    Seat utilization: 317/422 = 75.1% | 3-mo usage: 326 / 330 / 333 = +2.1% (Jun→Aug)
    RISK: LOW — Usage has grown every month for 12 months (289 → 333, +15.2% YoY) with 75.1% seat utilization.

14. C-0B144C78 | CSM: Cole Ingram | ARR $30,899.00
    Date used: 2026-11-02 (both systems agree; 12-mo single-year)
    Seat utilization: 169/224 = 75.4% | 3-mo usage: 101 / 101 / 106 = +5.0% (Jun→Aug)
    RISK: LOW — Usage is up +17.8% YoY (90 → 106) with 75.4% seat utilization.

15. C-0FC4DBB8 | CSM: Elena Sinclair | ARR $94,732.00
    Date used: 2026-11-05 (both systems agree; 12-mo single-year)
    Seat utilization: 356/464 = 76.7% | 3-mo usage: 189 / 191 / 193 = +2.1% (Jun→Aug)
    RISK: LOW — Largest ARR in the book is backed by steady growth (+14.9% YoY, 168 → 193) and 76.7% seat utilization.

16. C-0D5BBE3A | CSM: Dana Mercer | ARR $39,740.00
    Date used: 2026-11-09 (both systems agree; 12-mo single-year)
    Seat utilization: 85/102 = 83.3% | 3-mo usage: 88 / 90 / 91 = +3.4% (Jun→Aug)
    RISK: LOW — Usage up +19.7% YoY (76 → 91) with 83.3% seat utilization — strongest utilization/growth pairing in the book.

17. C-0FB9D5AF | CSM: Cole Ingram | ARR $63,158.00
    Date used: 2026-11-13 (both systems agree; 12-mo single-year)
    Seat utilization: 144/199 = 72.4% | 3-mo usage: 173 / 173 / 176 = +1.7% (Jun→Aug)
    RISK: LOW — Usage up +14.3% YoY (154 → 176) with 72.4% seat utilization.

18. C-0B344485 | CSM: Elena Sinclair | ARR $64,384.00
    Date used: 2026-11-16 (both systems agree; 12-mo single-year)
    Seat utilization: 224/287 = 78.0% | 3-mo usage: 238 / 240 / 244 = +2.5% (Jun→Aug)
    RISK: LOW — Usage up +15.6% YoY (211 → 244) with 78.0% seat utilization.

19. C-0CB2C1B4 | CSM: Dana Mercer | ARR $40,628.00
    Date used: 2026-11-20 (both systems agree; 12-mo single-year)
    Seat utilization: 386/473 = 81.6% | 3-mo usage: 47 / 48 / 49 = +4.3% (Jun→Aug)
    RISK: MEDIUM — Usage grew slightly (47 → 49) but only 49 of 386 seats-used were active in August (12.7%), the widest provisioning-engagement gap in the book.

20. C-22170CA1 | CSM: Cole Ingram | ARR $45,646.00
    Date used: 2026-11-24 (both systems agree; 12-mo single-year)
    Seat utilization: 251/294 = 85.4% | 3-mo usage: 143 / 148 / 146 = +2.1% (Jun→Aug)
    RISK: LOW — Usage up +12.3% YoY (130 → 146) with 85.4% seat utilization.

DATE DISAGREEMENTS — ALL FLAGGED (5 of 20)
1. C-0B7D2C30: ChurnZero 2026-09-10 vs Chargebee 2026-09-15 (5 days). 36-mo multi-year → trusted Chargebee.
2. C-0BCDB8C2: ChurnZero 2027-09-18 vs Chargebee 2026-09-18 (full year). 36-mo multi-year → trusted Chargebee.
3. C-0D2AB865: ChurnZero 2026-09-10 vs Chargebee 2026-09-22 (12 days). 24-mo multi-year → trusted Chargebee.
4. C-0BBE3E60: ChurnZero 2027-09-26 vs Chargebee 2026-09-26 (full year). 24-mo multi-year → trusted Chargebee.
5. C-0F5D2323: ChurnZero 2026-09-10 vs Chargebee 2026-09-29 (19 days). 24-mo multi-year → trusted Chargebee.
Remaining 15 accounts: dates identical across both systems — no disagreement.

DATA NOTE (not a date disagreement): in 5 accounts (C-0B20DB64 294 actives vs 214 seats used; C-0B369871 333 vs 317; C-0D5BBE3A 91 vs 85; C-0FB9D5AF 176 vs 144; C-0B344485 244 vs 224) August active_users exceeds seats_used, so seats_used behaves as a provisioning snapshot rather than a month-active count. No data provided to reconcile this.

TOTALS
Total ARR renewing (all 20, trusted dates 2026-09-15 → 2026-11-24): $1,048,715.00
  Arithmetic: 65,901 + 54,427 + 38,022 + 30,993 + 90,647 + 79,419 + 21,770 + 56,374 + 48,815 + 46,230 + 16,740 + 35,062 + 85,128 + 30,899 + 94,732 + 39,740 + 63,158 + 64,384 + 40,628 + 45,646 = 1,048,715

ARR AT RISK
  HIGH (6 accounts — C-0B7D2C30, C-0BCDB8C2, C-0D2AB865, C-0BBE3E60, C-0F5D2323, C-0EC6999D): $359,409.00 = 34.3% of book
    Arithmetic: 65,901 + 54,427 + 38,022 + 30,993 + 90,647 + 79,419 = 359,409
  MEDIUM (3 accounts — C-0BC34584, C-0B7A7546, C-0CB2C1B4): $92,430.00 = 8.8%
    Arithmetic: 16,740 + 35,062 + 40,628 = 92,430
  LOW (11 accounts): $596,876.00 = 56.9%
    Arithmetic: 1,048,715 − 359,409 − 92,430 = 596,876
  HIGH + MEDIUM exposure: $451,839.00 = 43.1% of book (359,409 + 92,430)

Notable concentration: 4 of the 5 date disagreements and all 3 past-due renewals sit in the multi-year cohort, and 5 of 6 HIGH-risk accounts renew on or before 2026-10-03 — $359,409 of at-risk ARR lands within the first 10 days of the window.
ticket-theme-synthesis0.80
Plain-English question: What are the real themes in this quarter's support tickets, ignoring the existing tags?
Prompt — what the model saw
Synthesize the quarter's support tickets into 4 to 7 themes. Do not trust existing tags; read the text. Per theme: name, count, share, distinct accounts, ARR affected, two ticket ids, one-line recommendation. Separate broad patterns from single-account noise and rank by ARR exposure, not volume.
cost $0.0112 · 105s · in 8,062 / out 8,783 tok
✗ theme_count_range✓ counts_verifiable✓ single_account_flagged✓ arr_ranked✓ ticket_ids

✓ Correct Answer (ground truth)

Plain-English answer: 80 tickets in 4-7 themes — points-not-posting leads (20 tickets, 25%, $31,100 ARR affected) — each with counts, distinct accounts, ARR, and a recommendation.
The answer in one line: total_tickets = 80; quarter_window = 2026-06-01 through 2026-08-31
total_tickets80
noteexisting tags are noisy/unreliable; classify by text
quarter_window2026-06-01 through 2026-08-31
ranking_by_arr_exposurehris_provisioning, redemption_failures, billing_dispute, points_not_posting, slack_sync
sample_ticket_ids80 items — e.g. IC-460004, IC-460016, IC-460001, IC-460041
single_account_noise_themebilling_dispute
themes{"points_not_posting": {"count": 20, "share_pct": 25.0, "distinct_accounts": 9, "arr_affected": "31100.00", "example_ticket_ids": ["IC-460001", "IC-460002"], "single_account": false}, "slack_sync": {"
graded checkstheme_count_range · counts_verifiable · single_account_flagged · arr_ranked · ticket_ids

Model's actual answer

show the model's full answer
SUPPORT TICKET THEME SYNTHESIS — Q3 window 2026-06-01 to 2026-08-29, 80 tickets, 19 distinct accounts
Themes derived from body_text only (existing tags ignored — they are unreliable, e.g. IC-460020 tagged "billing" is a points issue).

SHARE ARITHMETIC: count / 80 tickets. ARR affected = sum of distinct account ARR in theme (each account counted once).

=== BROAD PATTERNS (multi-account), RANKED BY ARR EXPOSURE ===

1. HRIS PROVISIONING FAILURES — new hires not created / skipped silently
   Count: 12/80 = 15.0% (4 "not creating accounts" + 3 "sync skipped 12 new hires" + 5 "not provisioned from HRIS sync")
   Distinct accounts: 3 — C-0B2213A9, C-0DDFC9A7, C-0F6C0F34
   ARR affected: 36,000 + 48,000 + 30,000 = $114,000
   Tickets: IC-460059, IC-460062
   Note: log shows no errors while hires are skipped — silent failure mode.
   Recommendation: Add reconciliation alert on HRIS sync (expected-vs-created headcount diff) before this hits renewals at C-0DDFC9A7 and C-0B2213A9.

2. REDEMPTION / GIFT CARD FAILURES — checkout hangs, codes never arrive, points deducted on error
   Count: 18/80 = 22.5% (6 "gift card email never showed up" + 5 "order errored but points deducted" + 4 "checkout spins" + 3 "code never arrived")
   Distinct accounts: 7 — C-0CEF69FD, C-0B827671, C-0FCCD2DF, C-0F876796, C-0D9CA315, C-0B0F1BAB, C-14264ABD
   ARR affected: 8,900 + 10,700 + 9,600 + 8,700 + 9,600 + 10,300 + 11,000 = $68,800
   Tickets: IC-460035, IC-460024
   Note: widest account spread of any theme — this is the true cross-book product defect.
   Recommendation: Fix the atomicity bug (points must not deduct when the gift card order fails) and add checkout retry with delivery confirmation.

3. POINTS LEDGER LAG — recognitions delivered but points never post (individual and team-wide)
   Count: 20/80 = 25.0% — the largest theme by volume (8 "delivered but points never arrived" + 5 "whole team after weekend" + 4 "last week's recognition" + 3 "balance not updated")
   Distinct accounts: 9 — C-0D3278C7, C-0BF20542, C-0D0B047C, C-0D284E42, C-0BE96399, C-0DD0626C, C-0B2895EF, C-0D6CC8E3, C-21FEBCBB
   ARR affected: 3,500 + 4,500 + 4,500 + 3,400 + 2,700 + 2,500 + 2,900 + 4,200 + 2,900 = $31,100
   Tickets: IC-460004, IC-460016
   Note: 9 accounts but all under $5k ARR — high volume, low ARR exposure. Also concentrated in C-0D3278C7 (4 tickets).
   Recommendation: Audit the weekend batch job behind "whole team" failures and add a delivered-vs-credited reconciliation metric.

4. SLACK INTEGRATION DEGRADATION — sync stops/toggle resets, auth won't stick, slash command errors
   Count: 14/80 = 17.5% (5 slash command errors + 4 "toggle resets itself" + 3 "stopped syncing" + 2 "re-auth does not stick")
   Distinct accounts: 4 — C-0BA71F12, C-10A56B0F, C-0B843542, C-8C2E8F00
   ARR affected: 3,900 + 5,400 + 4,400 + 5,200 = $18,900
   Tickets: IC-460047, IC-460046
   Note: C-0BA71F12 accounts for 7/14 tickets — partially account-concentrated, but all 4 accounts hit the slash-command failure, so the auth/token path is the broad root.
   Recommendation: Fix the token-refresh bug behind the self-resetting sync toggle and failing slash commands; this is one auth-layer defect, not four symptoms.

=== SINGLE-ACCOUNT NOISE (excluded from broad ranking) ===

5. BILLING / INVOICING ERRORS — seat-count discrepancies and wrong renewal tier price
   Count: 16/80 = 20.0% (6 wrong tier price + 5 "charged 200 seats, license 150" + 3 "third invoice in a row" + 2 "seat count never approved")
   Distinct accounts: 1 — C-0E9C27D1 (ALL 16 tickets)
   ARR affected: $52,000 (single account)
   Tickets: IC-460069, IC-460078
   Verdict: single-account noise despite 20% of ticket volume — one account in an active billing dispute across three invoice cycles. Do not treat as product work.
   Recommendation: Route to billing ops as one account escalation (correct the seat count to 150 and re-price the renewal tier), not a platform fix.

=== RANKING SUMMARY (ARR exposure, broad patterns only) ===
  1. HRIS provisioning        $114,000   12 tkts  3 accts
  2. Redemption/gift cards     $68,800   18 tkts  7 accts
  3. Points ledger lag         $31,100   20 tkts  9 accts
  4. Slack integration         $18,900   14 tkts  4 accts
  (noise) Billing/invoicing    $52,000   16 tkts  1 acct  — C-0E9C27D1 only

Volume-vs-ARR inversion to note: the biggest theme by count (points lag, 25%) is 3rd by ARR exposure ($31.1k), while the top ARR theme (HRIS, $114k) is only 15% of tickets. Ranking by volume would misdirect engineering away from the two themes that touch $182.8k of ARR combined.
marketing 5 tests
case-study-matchPASS 1.00
Plain-English question: Which three existing customers are the best case-study matches for this prospect?
Prompt — what the model saw
Given the prospect, rank the three most similar existing customers with public case studies for use as social proof. Similarity on industry, size band, use case, region. Cite the fields that drove each match. Do not name any customer without a case study and never include billing data or contact names.
cost $0.0027 · 24s · in 3,732 / out 1,146 tok
✓ top3✓ decoys_excluded✓ cites_fields✓ no_billing

✓ Correct Answer (ground truth)

Plain-English answer: Three ranked matches on industry, size, use case, and region (e.g. C-11C31562), with the higher-scoring decoy explicitly excluded for a mismatched field.
The answer in one line: prospect_alias = C-82AF3719
constraintsno_billing_data, no_contact_names, only_customers_with_has_case_study_true_named
decoys_excludedcompany_alias=C-D6217CAA; similarity_score=4; reason=no_case_study | company_alias=C-C153868F; similarity_score=4; reason=no_case_study
prospect_aliasC-82AF3719
prospect_fields_usedindustry, size_band, use_case, region
top3company_alias=C-11C31562; similarity_score=3 | company_alias=C-64171065; similarity_score=3 | company_alias=C-A13C193D; similarity_score=3
graded checkstop3 · decoys_excluded · cites_fields · no_billing

Model's actual answer

show the model's full answer
Prospect C-82AF3719: Technology / Mid-Market / employee_recognition / NA-West

Method: 1 point per exact field match across the 4 comparison fields (industry, size_band, use_case, region). Only has_case_study=true rows are eligible. Eligible pool: 8 customers.

Scoring of eligible candidates (matches / 4):
- C-64171065: industry ✓, size_band ✓, use_case ✓, region ✗ = 3
- C-11C31562: industry ✗, size_band ✓, use_case ✓, region ✓ = 3
- C-A13C193D: industry ✓, size_band ✓, use_case ✗ (retention), region ✓ = 3
- C-CD4829A7: industry ✓, size_band ✓, use_case ✗, region ✗ = 2
- C-9D69D147: industry ✗, size_band ✓, use_case ✗, region ✓ = 2
- C-FEFCC020: 0 | C-11209339: 0 | C-0CE315B2: 0

Three-way tie at 3/4. Tiebreak applied: use_case match first (strongest social-proof driver), then industry match.

Rank 1 — C-64171065 (3/4)
  Drove match: industry=Technology (exact), size_band=Mid-Market (exact), use_case=employee_recognition (exact).
  Miss: region=NA-East vs prospect NA-West.

Rank 2 — C-11C31562 (3/4)
  Drove match: use_case=employee_recognition (exact), region=NA-West (exact), size_band=Mid-Market (exact).
  Miss: industry=Manufacturing vs prospect Technology.

Rank 3 — C-A13C193D (3/4)
  Drove match: industry=Technology (exact), size_band=Mid-Market (exact), region=NA-West (exact).
  Miss: use_case=retention vs prospect employee_recognition.

Note: ranks 2 and 3 are separated only by the stated tiebreak (use_case over industry); both are 3/4 and the ordering between them is a judgment call, not a data distinction. No billing data or contact names referenced; no non-case-study customers named.
channel-attribution-mathPASS 1.00
Plain-English question: Compute cost per SQM, cost per SQO, and pipeline per dollar for every channel over six months.
Prompt — what the model saw
Compute channel performance for the trailing 6 months. Paid: spend, SQMs, SQOs, cost per SQM, cost per SQO, SQM-to-SQO rate, pipeline amount, pipeline per dollar. Organic: volume, SQO rate, pipeline. Report a channel with spend and zero SQMs as undefined, not zero. Flag rows where SQO date precedes SQM date. Finish with a reallocation recommendation and your confidence given sample sizes.
cost $0.0124 · 98s · in 13,255 / out 7,337 tok
✓ undefined_channel✓ undefined_not_zero✓ flagged_rows✓ best_channel✓ reallocation

✓ Correct Answer (ground truth)

Plain-English answer: LinkedIn Ads at $960/SQM and $3,000/SQO, organic search at $90,000 pipeline on zero spend, paid social flagged at $18,000 with zero SQMs, and a channel winner named.
The answer in one line: best_pipeline_per_dollar_channel = paid_search
best_pipeline_per_dollar_channelpaid_search
confidence_notewebinars and referral have small SQM samples (<20); treat rates as directional.
flagged_rows_sqo_before_sqmCT-000041, CT-000044
months2026-03, 2026-04, 2026-05, 2026-06, 2026-07, 2026-08
per_channel{"linkedin_ads": {"spend_usd": 24000, "sqms": 25, "sqos": 8, "pipeline_amount": 96000, "cost_per_sqm": 960.0, "cost_per_sqo": 3000.0, "sqm_to_sqo_rate": 0.32, "pipeline_per_dollar": 4.0}, "organic_sea
reallocation_recommendationfrom=paid_social; to=paid_search; rationale=paid_social has spend but zero SQMs (cost per SQM undefined, not zero); paid_search has the highest pipeline per dollar.
undefined_channelspaid_social
graded checksundefined_channel · undefined_not_zero · flagged_rows · best_channel · reallocation

Model's actual answer

show the model's full answer
Window = 2026-03 through 2026-08 (the full data window; all 122 contacts fall in these 6 months).

PAID CHANNELS
spend from channel_spend.csv (6 months x monthly rate):
  paid_search  6 x $6,000 = $36,000
  linkedin_ads 6 x $4,000 = $24,000
  paid_social  6 x $3,000 = $18,000
  webinars     6 x $1,500 =  $9,000
  TOTAL PAID              = $87,000

channel      spend    SQMs  SQOs  cost/SQM        cost/SQO        SQM->SQO    pipeline    pipeline/$
paid_search  $36,000   40    18   36000/40=$900   36000/18=$2,000   18/40=45.0%  $720,000   720000/36000=$20.00
linkedin_ads $24,000   25     8   24000/25=$960   24000/8=$3,000     8/25=32.0%   $96,000    96000/24000=$4.00
paid_social  $18,000    0     0   UNDEFINED       UNDEFINED        UNDEFINED        $0.00   UNDEFINED (spend, zero SQMs)
webinars     $ 9,000   12     5   9000/12=$750    9000/5=$1,800     5/12=41.7%   $60,000    60000/9000=$6.67

  paid_social: $18,000 spend, 0 SQMs -> cost per SQM / cost per SQO / SQM-to-SQO rate / pipeline per dollar are all UNDEFINED (not zero). Pipeline amount is $0 (observed, since no records exist).

Paid totals: 77 SQMs, 31 SQOs, blended SQM->SQO 31/77 = 40.3%, pipeline $876,000, pipeline/$ = 876000/87000 = $10.07.

ORGANIC CHANNELS (no spend data provided)
channel         volume (SQMs)  SQOs  SQO rate      pipeline
organic_search      30          10   10/30=33.3%   $90,000
referral            15           6    6/15=40.0%   $48,000
TOTAL               45          16   16/45=35.6%  $138,000

DATA QUALITY FLAG — SQO date precedes SQM date:
  CT-000044, linkedin_ads: SQM 2026-07-23, SQO 2026-07-18 (5 days before), pipeline $12,000
  CT-000041, linkedin_ads: SQM 2026-06-14, SQO 2026-06-09 (5 days before), pipeline $12,000
  Both in linkedin_ads. Excluding them: linkedin_ads SQOs 8 -> 6, pipeline $96,000 -> $72,000,
  cost/SQO $3,000 -> $4,000, SQM->SQO 32.0% -> 24.0%, pipeline/$ $4.00 -> $3.00.
  No other rows inverted; no rows carry pipeline without an SQO date.

REALLOCATION RECOMMENDATION
1. Cut paid_social entirely ($18,000 over 6 months, 0 SQMs — it produced no measurable funnel entry).
   Reallocate to paid_search and webinars, the two channels with the best cost per SQO
   ($2,000 and $1,800) and best pipeline per dollar ($20.00 and $6.67).
   Suggested split: ~$12,000 to paid_search, ~$6,000 to webinars (weighted by paid_search's
   higher absolute SQO volume 18 vs 5 and higher pipeline/$).
2. Consider trimming linkedin_ads: highest cost per SQO ($3,000; $4,000 once the two inverted-date
   rows are excluded) and lowest pipeline/$ ($4.00) among spenders, and 2 of its 8 SQOs carry
   date-integrity defects. Trim to a test budget and hold the rest pending cleaner attribution.

CONFIDENCE
- paid_search: moderate-to-high. 40 SQMs / 18 SQOs is the largest paid sample. Caveat: every
  paid_search pipeline amount is exactly $40,000 — uniform values suggest a default/staged amount,
  so the $720,000 and $20.00 pipeline/$ rest on unverified per-deal values.
- webinars: low-to-moderate. 5 SQOs and 12 SQMs — cost/SQO is volatile (one SQO swings it ~$450).
- linkedin_ads: low-to-moderate. 8 SQOs (6 clean); results move materially on the 2 defective rows.
- paid_social: the "zero SQMs" finding is high-confidence for this dataset (0 of 122 contacts), but
  a 6-month zero may reflect tracking gaps as easily as true non-performance — verify attribution
  before permanently cutting spend.
- organic (30 and 15 SQMs): moderate; rates are directional only since no spend/cost basis exists.
Overall: directionally sound to move paid_social budget; exact split should be treated as a test
hypothesis, not a forecast.
competitor-battlecard-updatePASS 1.00
Plain-English question: Update the Rivally battlecard with sourced pricing and where each side wins.
Prompt — what the model saw
Update the battlecard for the competitor. Sections: one-line positioning, pricing with source and date (newer source wins, note the conflict), where they win, where we win, objections and responses, recent changes, our 12-month win/loss record against them. Cite a snippet id for every factual claim. Rep opinion on a call is not a fact about the competitor. Anything from the old card you cannot re-source gets marked unverified.
cost $0.0040 · 40s · in 4,554 / out 2,284 tok
✓ competitor_named✓ pricing✓ win_loss✓ unverified_marked✓ citations

✓ Correct Answer (ground truth)

Plain-English answer: Rivally at $7 per user per month (the August source wins over the stale $5, conflict noted), a 13-7 win/loss record, and two unverified old claims flagged.
The answer in one line: competitor = Rivally
citation_ids_available25 items — e.g. S01, S02, S03, S04
competitorRivally
pricingvalue_per_user_per_month=7; source_snippet=S17; source_date=2026-08-12
rep_opinion_snippet_ids_not_factsS09, S21
unverified_old_claimsRivally lacks a Slack integration., Rivally was acquired by WorkHuman in 2025.
win_loss_12mowins=13; losses=7
graded checkscompetitor_named · pricing · win_loss · unverified_marked · citations

Model's actual answer

show the model's full answer
# Battlecard: Rivally — Updated 2026-09-25

## One-line positioning
Points-based employee recognition for mid-market, with an engaging recognition feed but thin analytics and admin tooling (S02, S16, S07).

## Pricing
- Current list: $7 per user/month, Recognition Starter, annual billing required (S17, pricing_page, 2026-08-12).
- CONFLICT / price history: $5/user/mo on 2026-01-20 (S03) and still $5 on 2026-04-01 (S08). Newer source (S17) wins — the old card's "$5 as of 2026-01" is stale.
- Deal quotes: $6.50/user/mo to a 500-seat prospect, annual term (S13, 2026-06-02); $7/user/mo list with 15% off for a 3-year term = $5.95/user/mo effective (S18, 2026-08-14; arithmetic: 7 × 0.85 = 5.95).
- Rivally Pulse survey add-on is priced separately, not bundled (S23, 2026-09-01).
- Excluded as non-fact: "discounting aggressively" is rep opinion only (S21).

## Where they win
- Distributed EU teams / multi-language (S12, g2_review, 2026-05-21).
- EU data residency — GA with Dublin office (S15, 2026-07-01); pitched to prospects earlier (S05).
- Fast time-to-value: setup under a week, Slack integration works out of the box (S04, 2026-02-02).
- Support responsiveness: under 4 hours (S22, 2026-08-30).
- Engaging, points-based recognition feed (S02, S16).

## Where we win
- Analytics depth: reporting dashboards are "basic compared to enterprise tools" (S07, 2026-03-22); an 800-seat prospect chose Bonusly over Rivally citing analytics depth (S25, 2026-09-03).
- Admin/IT readiness: no SCIM provisioning, manual user management (S10, 2026-04-10); admin tooling lags peers (S16); no bulk recognition editing (S24, 2026-09-02).
- Migration/exit: CSV-only analytics exports made offboarding hard (S20, 2026-08-25).
- US rewards catalog depth: EMEA catalog is thinner than US (S14, 2026-06-02).

## Objections and responses
1. "Rivally is $5/user/mo." → Out of date. Their own pricing page shows $7 as of 2026-08-12 (S17); quotes in-market landed at $6.50–$7 (S13, S18).
2. "Rivally has EU data residency and multi-language." → True (S15, S12) — don't contest it. Pivot to analytics and admin tooling gaps (S07, S10, S24), and note the thinner EMEA rewards catalog (S14).
3. "Setup is fast and Slack works out of the box." → True (S04) — acknowledge, then pivot to what happens after setup: basic reporting (S07), no SCIM (S10), CSV-only exports on exit (S20).
4. "Rivally is discounting aggressively." → No evidence. Only a rep opinion (S21), not usable as a claim. Known data point is 15% off list for a 3-year term (S18).

## Recent changes (last ~6 months)
- 2026-09-01: Rivally Pulse exits beta; priced as add-on, not bundled (S23).
- 2026-08-12: list price up from $5 to $7/user/mo (S17).
- 2026-08-20: Microsoft Teams app v2 in public preview (S19).
- 2026-07-01: Dublin office opens; EU data residency GA (S15).
- 2026-05-09: hires ex-Workday VP EMEA to lead European expansion (S11).
- 2026-03-05: Rivally Pulse engagement survey add-on launched (S06).

## 12-month win/loss vs. Rivally (2025-09 through 2026-08)
- Wins: 13 — Deal-A9FD43, Deal-7AA785, Deal-44C524, Deal-0D0CD6, Deal-D5B790, Deal-5C636E, Deal-67BE14, Deal-1B6969, Deal-F03E7B, Deal-072E31, Deal-F65C8F, Deal-E46EAB, Deal-1D2392
- Losses: 7 — Deal-7767F5, Deal-5645A5, Deal-C6FFAA, Deal-D263E0, Deal-935746, Deal-9066A6, Deal-72A02F
- Record: 13–7 across 20 deals. Win rate = 13 ÷ 20 = 65%.
- Monthly: 2025-09 2–0... correcting per-file: 2025-09 (1W–1L), 2025-10 (2W–0L), 2025-11 (1W–1L), 2025-12 (1W–1L), 2026-01 (2W–0L), 2026-02 (2W–0L), 2026-03 (1W–1L), 2026-04 (0W–2L), 2026-05 (0W–1L), 2026-06 (1W–0L), 2026-07 (1W–0L), 2026-08 (1W–0L). Sums check: 13W + 7L = 20.

## Old card claims — disposition
- "Pricing starts at $5/user/mo (as of 2026-01)" → SUPERSEDED. Was accurate at the time (S03) but now $7 (S17).
- "Rivally lacks a Slack integration" → FALSE. Contradicted by S04 (Slack integration worked out of the box, 2026-02-02). Remove.
- "Rivally was acquired by WorkHuman in 2025" → UNVERIFIED. No snippet in the provided corpus supports this. Do not use.
- "Strong in EU enterprise with multi-language support" → RE-SOURCED to S12.
- "Points-based recognition for mid-market" → RE-SOURCED to S02 (points-based feed; mid-market and enterprise reviewers both present in corpus, so "mid-market" specifically is only partially supported — S04 is a mid-market reviewer).
nurture-sequence-diagnosisPASS 1.00
Plain-English question: What is broken in each email sequence — rates, overlaps, tracking errors?
Prompt — what the model saw
Review the sequences. Per sequence: sent, open rate, reply rate, meeting rate, weakest step. Identify tracking errors (opened above sent), audience overlap between sequences, and the failure mode for anything under 2 percent reply. One change per weak sequence and which to fix first. Under 350 words.
cost $0.0216 · 439s · in 28,408 / out 10,289 tok
✓ tracking_error✓ overlap_pair✓ weakest_step✓ fix_priority

✓ Correct Answer (ground truth)

Plain-English answer: Per-sequence open, reply, and meeting rates with the weakest step (all three weakest at step 3), an opened-above-sent tracking error, and audience overlap flagged.
The answer in one line: fix_priority_first = Cold Outbound - HR Leaders
all_pairwise_overlaps['Cold Outbound - HR Leaders', 'Cold Outbound - People Ops'], ['Expansion Nurture', 'New Logo Nurture']
audience_overlap_pairCold Outbound - HR Leaders, Cold Outbound - People Ops
failure_modeCold Outbound - HR Leaders=opens healthy (~37%) but replies near zero from step 1 -> body/CTA not landing, not a deliverability or subject-line problem
fix_priority_firstCold Outbound - HR Leaders
per_sequence{"New Logo Nurture": {"sent": 1386, "opened": 490, "replied": 90, "meetings": 27, "open_rate": 0.3535, "reply_rate": 0.0649, "meeting_rate": 0.0195, "weakest_step": 3}, "Expansion Nurture": {"sent": 8
tracking_errorsequence=Expansion Nurture; step=2; sent=300; opened=340; issue=opened exceeds sent
under_2pct_replyCold Outbound - HR Leaders
graded checkstracking_error · overlap_pair · weakest_step · fix_priority

Model's actual answer

show the model's full answer
Per-sequence rates (rates = count / sent for that step):

New Logo Nurture — sent 500/458/428. Open 42.0%/34.9%/28.0% (210/500, 160/458, 120/428). Reply 8.40%/6.55%/4.21%. Meeting 2.40%/1.97%/1.40%. Weakest step: 3. Totals: sent 1,386, open 35.4%, reply 6.49%, meeting 1.95%.

Expansion Nurture — sent 300/300/275. Open 43.3%/113.3% (340/300 — invalid)/34.5%. Reply 7.33%/8.33%/4.36%. Meeting 1.67%/1.33%/1.09%. Weakest step: 3. Totals (inflated by the error): sent 875, open 64.6%, reply 6.74%, meeting 1.37%.

Cold Outbound - HR Leaders — sent 600/595/590. Open 40.0%/29.4%/22.0%. Reply 0.83%/0.34%/0.17%. Meeting 0.00% all steps. Weakest step: 3. Totals: sent 1,785, open 30.5%, reply 0.45%, meeting 0.00%.

Cold Outbound - People Ops — sent 400/386/377. Open 37.5%/28.5%/21.2%. Reply 3.50%/2.33%/1.59%. Meeting 0.75%/0.52%/0.27%. Weakest step: 3. Totals: sent 1,163, open 29.2%, reply 2.49%, meeting 0.52%.

Tracking errors: Expansion Nurture step 2 — opened 340 > sent 300 (113%). Only such row; also makes Expansion's aggregate open rate (64.6%) unusable.

Audience overlap (963 roster rows, deduped on contact_key): 23 contacts appear in two sequences. 21 pairs cross Cold Outbound - HR Leaders and Cold Outbound - People Ops (e.g. CT-000849, CT-001105, CT-001130, CT-001345). 2 cross Expansion Nurture and New Logo Nurture (CT-000301, CT-000624). Roster counts (299/157/290/217) are far below step-1 sent (500/300/600/400), so the roster appears partial — contact-level overlap may be understated.

Under 2% reply — failure mode: Cold Outbound - HR Leaders (all 3 steps) and Cold Outbound - People Ops step 3. Opens are healthy (22–40%), so this is not deliverability — it's a messaging/CTA failure: subject lines earn opens, body copy (generic cold pitch to exec titles) earns no response. HR Leaders never converts at all (0 meetings on 1,785 sends).

One change each:
- Cold Outbound - HR Leaders: rewrite step 1 body to one specific, role-relevant ask (15-min call tied to an HR metric) — replace the generic pitch.
- Cold Outbound - People Ops: cut step 3 or replace with a breakup/one-line question; 1.59% reply (6/377) doesn't justify the send.

Fix first: Cold Outbound - HR Leaders — largest volume (1,785 sent), worst reply (0.45%), zero meetings.
weekly-marketing-goals-rollupPASS 1.00
Plain-English question: Are we on pace on the quarter's marketing goals?
Prompt — what the model saw
Produce the weekly marketing goals update for the active quarter: SQMs, SQOs, DS2s, closed-lost MIA rate, same-quarter close count, active pipeline coverage against target. For each: QTD actual, target, delta, pace (on/behind/ahead given days elapsed). One paragraph on what moved this week.
cost $0.0052 · 26s · in 9,301 / out 1,202 tok
✓ metrics_present✓ deltas✓ coverage_ratio✓ pace✓ paragraph

✓ Correct Answer (ground truth)

Plain-English answer: SQMs ahead of pace, SQOs on pace, DS2s behind (35 short of expected at day 66 of 92), with deltas and same-quarter closes.
The answer in one line: days_in_quarter = 92; days_elapsed = 66; coverage_ratio_active_pipeline = 0.75; elapsed_fraction = 0.7174 (+1 more below)
days_in_quarter92
coverage_ratio_active_pipeline0.75
days_elapsed66
elapsed_fraction0.7174
metrics{"SQMs": {"qtd_actual": 230, "target": 300, "delta": -70, "expected_at_day_elapsed": 215.22, "pace": "ahead"}, "SQOs": {"qtd_actual": 84, "target": 120, "delta": -36, "expected_at_day_elapsed": 86.09,
pace_ruleexpected = target * days_elapsed/days_in_quarter; within +/-5% of expected = on, above = ahead, below = behind; lower_better metrics invert.
this_week_paragraph_required1
graded checksmetrics_present · deltas · coverage_ratio · pace · paragraph

Model's actual answer

show the model's full answer
Q3-2026 MARKETING GOALS UPDATE — QTD as of day 66 of 92 (71.7% of quarter elapsed: 66/92)
Source: marketing_qtd.csv, targets.csv, quarter_meta.csv only.

METRIC                  QTD ACTUAL      TARGET      DELTA           PACE (vs linear 71.7% run-rate)
SQMs                    230             300         -70             AHEAD — pace run-rate = 300 x 66/92 = 215.2; 230 > 215.2 (+14.8)
SQOs                    84              120         -36             BEHIND — 120 x 66/92 = 86.1; 84 < 86.1 (-2.1)
DS2s                    40              75          -35             BEHIND — 75 x 66/92 = 53.8; 40 < 53.8 (-13.8)
Closed-lost MIA rate    20.0%           10%         +10 pts over    BEHIND — lower is better; 5/25 = 0.20 vs 0.10 target
Same-quarter closes     10              20          -10             BEHIND — 20 x 66/92 = 14.3; 10 < 14.3 (-4.3)
Active pipeline         $3,000,000      $4,000,000  -$1,000,000     AHEAD of linear, below full target — 4.0M x 66/92 = $2.87M; $3.0M > $2.87M (+$0.13M); coverage vs target = 3.0/4.0 = 0.75x

Arithmetic shown:
- MIA rate = closed_lost_mia / closed_lost_total = 5/25 = 0.20
- Pace run-rate for higher_better metrics = target x 66/92 = target x 0.7174
- Coverage ratio = active_pipeline / target = 3,000,000 / 4,000,000 = 0.75x

What moved this week: NOT AVAILABLE — the provided files contain a single QTD snapshot (marketing_qtd.csv) with no prior-week values or week-over-week deltas, so I cannot state what changed this week without inventing data. To produce this paragraph, provide last week's snapshot of the same six metrics.
communication 5 tests
ceo-slack-compressionPASS 1.00
Plain-English question: Compress the forecast analysis into one Slack message the CEO can read in ten seconds.
Prompt — what the model saw
Compress the forecast analysis below into a single Slack message for the CEO. Rules: at most three sentences; under 80 words; lead with the decision or the so-what; keep every distinct number that matters and the one caveat that changes the conclusion; cut any sentence that does not carry a claim, a number, or a directive. No em dashes, no throat-clearing phrases.
cost $0.0029 · 56s · in 3,675 / out 1,428 tok
✓ max_3_sentences✓ under_80_words✓ weighted_present✓ caveat_present✓ no_em_dash

✓ Correct Answer (ground truth)

Plain-English answer: Three sentences, under 80 words, decision first, every distinct number kept plus the one caveat — all checked by deterministic text rules.
The answer in one line: weighted_number = 115976; must_keep = data-quality caveat — deleting the bad news fails
weighted_number115,976
limitmax 3 sentences, max 80 words, no em-dashes
must_keepdata-quality caveat — deleting the bad news fails
graded checksmax_3_sentences · under_80_words · weighted_present · caveat_present · no_em_dash

Model's actual answer

show the model's full answer
Q3 forecast is 115,976.75 weighted (44,729 COMMIT at 100% plus 71,247.75 from 203,565 BEST_CASE at 35%, PIPELINE at 0) but do not rely on it. Why-buys are empty on all 7 COMMIT deals, so the 44,729 has no documented reason to buy behind it. 32 deals worth 227,575 close after 2026-09-30 and are excluded, including COMMIT Deal-D348E1 at 13,770.
stalled-deal-followup-emailPASS 1.00
Plain-English question: Write a follow-up email for this stalled deal — under 120 words, one prior touch, one ask.
Prompt — what the model saw
Write a follow-up email for this stalled deal. Under 120 words. Reference one specific prior touch from the engagement history, make one clear ask, and match the tone of the prior thread. No invented details.
cost $0.0020 · 22s · in 3,444 / out 560 tok
✓ word_cap✓ references_touch✓ single_ask

✓ Correct Answer (ground truth)

Plain-English answer: An email that references the July 28 milestone-awards demo, makes one clear ask, and invents nothing — Deal-0D2F7A, 31 days since last engagement.
The answer in one line: amount = 5100; word_cap = 120; days_since_last_engagement = 31; deal_alias = Deal-0D2F7A (+4 more below)
amount5,100
days_since_last_engagement31
word_cap120
allowed_facts14 items — e.g. Deal-0D2F7A, $5,100, DS3, Alex Franklin
allowed_newENG-00001, ENG-00002, ENG-00003, ENG-00004, ENG-00005
deal_aliasDeal-0D2F7A
must_referenceengagement_id=ENG-00002; date=2026-07-28; type=meeting; why=most substantive prior touch; concrete prospect reactions to milestone awards and points catalog
ownerAlex Franklin
single_cta1
stageDS3
tonewarm, concise, professional (matches prior thread)
graded checksword_cap · references_touch · single_ask

Model's actual answer

show the model's full answer
Subject: Following up — 150-seat pricing recap

Hi there,

Following up on my August 5 recap of our July 28 demo with your People team, which included pricing for 150 seats. I know priorities shift, so I want to keep this easy.

You'd asked for pricing after seeing the automated milestone awards and the points catalog — I'd like to make sure nothing's outstanding on your end.

Could we book 15 minutes to walk through any open questions on the numbers?

Best,
Alex Franklin

Word count (body + sign-off): 79 words — under the 120 limit. One prior touch cited (the Aug 5 recap email of the July 28 demo), one ask (15 minutes to walk through pricing questions). Tone matches the thread's concise, professional, recap-style outbound emails. No contact names, company name, or details beyond the CSVs were used.
funnel-digest-paragraphPASS 1.00
Plain-English question: Write the weekly GTM digest — one paragraph per team, every number traceable.
Prompt — what the model saw
Write the weekly GTM digest: one paragraph per team (Marketing, Sales, CS, PLG) in Bonusly brand voice. Every number must trace to the extracts. No customer contact names, no employee performance commentary, no compensation. Four paragraphs, under 400 words total.
cost $0.0025 · 32s · in 3,273 / out 1,196 tok
✓ four_paragraphs✓ word_cap✓ teams

✓ Correct Answer (ground truth)

Plain-English answer: Four paragraphs (Marketing, Sales, CS, PLG) using only the allowed numbers, no contact names, no performance commentary, under 400 words.
The answer in one line: word_cap = 400
word_cap400
allowed_numbers16 items — e.g. 2, 3, 9, 08
bannedcustomer contact names, employee performance commentary, compensation
required_paragraphsMarketing, Sales, CS, PLG
graded checksfour_paragraphs · word_cap · teams

Model's actual answer

show the model's full answer
Marketing closed the week at 46 SQMs against a 52 target — 88.5% to goal (46 ÷ 52 = 0.885) — with webinar carrying 18 of those SQMs (18 ÷ 46 = 39% of the total) and paid search cost per SQM holding steady at $150; the takeaway is a dependable, efficient base with webinar as the workhorse channel.

Sales converted 14 SQOs and set 9 DS2 meetings, generating $310,000 in new pipeline and 3 same-quarter closes — a week where top-of-funnel conversion translated directly into near-term revenue.

CS saved 2 renewals and moved team NPS to 61, heading into next week with 3 open red-flag accounts to protect; the save count and NPS trend are constructive, but those 3 red flags are where next week's attention will matter most.

PLG added 412 new signups at 31% activation, which works out to roughly 128 activated accounts (412 × 0.31 = 127.7), with 38 companies reaching the aha moment of 10 recognition gives — meaningful top-of-funnel volume with a clear milestone that's pulling a solid slice of signups into habit-forming usage.
partner-digest-sectionPASS 1.00
Plain-English question: Write the partner digest so every partner in the roster appears exactly once.
Prompt — what the model saw
Write the partner digest for the period. Every partner in the roster appears once: active plays get status with numbers, quiet partners get one line. Partner-sourced pipeline count and amount must match the deal data. Nothing omitted, nothing added.
cost $0.0022 · 16s · in 3,487 / out 754 tok
✓ all_partners✓ pipeline_amount✓ quiet_one_line

✓ Correct Answer (ground truth)

Plain-English answer: Four active partners with status and numbers, four quiet partners with one line each, and partner-sourced pipeline that reconciles to the deal data.
The answer in one line: rule = every roster partner named exactly once; quiet partners get one line; no partners outside roster.
active_no_dealsWorkWell Group
partner_sourced_pipelinecount=5; amount=550000
per_partner_pipeline{"Apex Rewards Co": {"count": 2, "amount": 275000}, "HRCloud Partners": {"count": 1, "amount": 140000}, "CultureBridge": {"count": 2, "amount": 135000}}
quiet_partnersRecogniq, KudosWave, PeopleFirst Advisors, TotalPerk
rosterpartner_name=Apex Rewards Co; status=active | partner_name=HRCloud Partners; status=active | partner_name=CultureBridge; status=active | partner_name=WorkWell Group; status=active | partner_name=Recogniq; status=quiet | partner_name=KudosWave; status=quiet … (+2 more)
ruleevery roster partner named exactly once; quiet partners get one line; no partners outside roster.
graded checksall_partners · pipeline_amount · quiet_one_line

Model's actual answer

show the model's full answer
Partner Digest — 2026-08-24 to 2026-09-04

ACTIVE PARTNERS

Apex Rewards Co — active
- Co-webinar locked for 09-15 (Dana Mercer, 2026-08-25).
- 2 sourced opps this period, both DS1, both logged with UTM Source = Partner (Cole Ingram confirmed).
- Deals: Deal-DDAAF2 ($180,000), Deal-2CF33E ($95,000).
- Pipeline: 2 deals, $275,000 (180,000 + 95,000 = 275,000).

HRCloud Partners — active
- Integration referral closed security review; opp moved to DS2 (Alex Franklin, 2026-08-27).
- 1 sourced opp this period.
- Deal: Deal-F1CDA5 ($140,000).
- Pipeline: 1 deal, $140,000.

CultureBridge — active
- Lunch-and-learn produced 2 sourced opps, both early stage (Farid Osman, 2026-08-29).
- Deals: Deal-096E1D ($60,000), Deal-067213 ($75,000).
- Pipeline: 2 deals, $135,000 (60,000 + 75,000 = 135,000).

WorkWell Group — active
- No sourced deals this period; joint playbook restart planned for Q4, planning call booked 09-09 (Dana Mercer, 2026-09-02).
- Pipeline: 0 deals, $0.

QUIET PARTNERS

- Recogniq — no activity this period.
- KudosWave — no activity this period.
- PeopleFirst Advisors — no activity this period.
- TotalPerk — no activity this period.

TOTALS (all partner-sourced deals in partner_deals.csv, UTM Source = Partner)
- Deal count: 5 (Apex 2 + HRCloud 1 + CultureBridge 2)
- Amount: $550,000 (275,000 + 140,000 + 135,000 + 0 = 550,000)

Note: all 5 deals carry UTM Source = Partner in the deal data, consistent with the Slack-logged sourced opps (Apex 2, HRCloud 1, CultureBridge 2). Stage labels (DS1/DS2/early) come from Slack only; the deal file does not include stage fields, so they are reported as stated, not verified against deal data.
rfp-answer-from-policyPASS 1.00
Plain-English question: Answer the 10 RFP items from the policy excerpts only — no guessing.
Prompt — what the model saw
Answer the 10 questionnaire items from the policy excerpts only. Cite the excerpt id for every answer. Where the excerpts do not answer the question, mark it unanswerable and say what document would; do not guess. Tone: vendor response, direct.
cost $0.0020 · 14s · in 3,450 / out 536 tok
✓ answered_count✓ unanswerable_flagged✓ citations✓ no_guess

✓ Correct Answer (ground truth)

Plain-English answer: Seven answered with excerpt citations (Q1-Q7), three flagged unanswerable with the document that would answer them (Q8-Q10).
The answer in one line: answered_count = 7; rule = unanswerable items must be flagged and the document that would answer them named; do not guess.
answered_count7
answer_map{"Q1": ["P1"], "Q2": ["P1"], "Q3": ["P2"], "Q4": ["P3"], "Q5": ["P4"], "Q6": ["P5"], "Q7": ["P6"]}
ruleunanswerable items must be flagged and the document that would answer them named; do not guess.
unanswerableQ8, Q9, Q10
valid_citation_idsP1, P2, P3, P4, P5, P6, P7, P8
graded checksanswered_count · unanswerable_flagged · citations · no_guess

Model's actual answer

show the model's full answer
Responses to questionnaire (source: policy_excerpts.md only)

Q1 — Is customer data encrypted at rest? If so, how?
Yes. All customer data is encrypted at rest using AES-256. [P1]

Q2 — Is data encrypted in transit?
Yes. Data in transit is encrypted with TLS 1.2 or higher. [P1]

Q3 — Do you hold a SOC 2 report, and can we obtain a copy?
Yes. Bonusly maintains a SOC 2 Type II report, renewed annually. It is available to prospects under NDA. [P2]

Q4 — How long is customer data retained after contract termination?
90 days after contract termination, then deleted from production systems. (Retention = duration of contract + 90 days; the post-termination figure is the 90-day tail.) [P3]
Note: P3 states deletion from production systems only; the excerpts do not address backup deletion timelines (backups retained 35 days per [P6], but the excerpts do not link the two).

Q5 — Where is your subprocessor list published and how are changes notified?
Published at trust.bonusly.com. Customers receive 30 days notice before a new subprocessor is added. [P4]

Q6 — Do you support SAML single sign-on?
Yes. SAML 2.0 single sign-on is available on the Enterprise plan. (SCIM user provisioning is also available on Enterprise.) [P5]

Q7 — What is your production backup schedule and retention?
Daily backups of production databases; backups are encrypted and retained for 35 days. [P6]

Q8 — Do you hold an ISO 27001 certification?
Unanswerable from the excerpts. The excerpts cover SOC 2 Type II only [P2]; no excerpt mentions ISO 27001. Needed: our certifications page at trust.bonusly.com or a current ISO 27001 certificate.

Q9 — What is your contractual uptime SLA percentage?
Unanswerable from the excerpts. No excerpt states an uptime SLA. [P7] covers incident notification only. Needed: the Master Subscription Agreement / SLA schedule in the contract.

Q10 — Will you sign a HIPAA Business Associate Agreement?
Unanswerable from the excerpts. No excerpt addresses HIPAA or BAAs. Needed: our legal/compliance position on BAAs (typically in the MSA or a compliance addendum).
ops-maintenance 5 tests
skill-file-reconciliationPASS 1.00
Plain-English question: Audit a skill set against its manifest and flag every conflict with a severity and an action.
Prompt — what the model saw
Reconcile the skill set below against its manifest. Find and report (with severity CRITICAL/WARNING/INFO and an action type TRIM_DESC, UPDATE_BODY, DELETE_SKILL, MERGE, or REVIEW): (1) any two skills whose ALWAYS-trigger phrases overlap/duplicate, (2) any circular delegation chain between skills (name it), (3) any delegation target that does not exist (dangling, name it), (4) any version conflict and which skill should survive, (5) how many manifest descriptions exceed 1,024 characters, (6) any hardcoded page ids, dates, or person names in skill bodies, (7) manifest drift in BOTH directions (files with no manifest row, and manifest rows with no file). Make one proposal per finding; rewrite nothing.
cost $0.0547 · 439s · in 85,336 / out 19,750 tok
✓ duplicate_trigger_found✓ circular_chain✓ dangling_target✓ oversized_desc_count✓ hardcoded_values✓ manifest_drift✓ severity_and_action

✓ Correct Answer (ground truth)

Plain-English answer: A severity-ranked list of TRIM_DESC / UPDATE_BODY / DELETE_SKILL / MERGE / REVIEW actions covering every rule violation found.
The answer in one line: files_audited = 14; rank = CRITICAL / WARNING / INFO with one fix action per issue
files_audited14
findingsduplicate trigger, circular delegation chain, dangling file reference, oversized description, hardcoded values, manifest drift
rankCRITICAL / WARNING / INFO with one fix action per issue
graded checksduplicate_trigger_found · circular_chain · dangling_target · oversized_desc_count · hardcoded_values · manifest_drift · severity_and_action

Model's actual answer

show the model's full answer
SKILL MANIFEST RECONCILIATION — 14 manifest rows vs 14 provided SKILL.md files
Scope: only the data provided (skill_manifest.csv + 14 bodies). Targets referenced outside this set are unverifiable from provided data and flagged as dangling "relative to this set."

================================================================
(1) ALWAYS-TRIGGER OVERLAP / DUPLICATION — 3 pairs
================================================================
1a. CRITICAL | MERGE | "comms-drafter" + "email-drafter"
Evidence (verbatim duplicate trigger phrases, both descriptions):
  "write me an email" · "draft a follow-up" · "help me reply" ·
  "what should I say" · "bump email" · "contract nudge"
comms-drafter adds "Trigger for: ..." list; email-drafter adds "review this email,"
"rewrite this." Neither disambiguates the other (both only hand off to
deal-strategy-coach). Same utterance routes to two skills with different output
contracts (comms-drafter = all external comms + support/partner; email-drafter =
email only + Gmail signature retrieval block).
Proposal: MERGE email-drafter into comms-drafter (survivor: comms-drafter, the
superset scope); preserve email-drafter's signature-retrieval block as a section.

1b. CRITICAL | TRIM_DESC | "weekly-pipeline-report" vs "pipeline-intelligence-report"
Overlapping trigger phrases: "run the pipeline report" / "do the pipeline report" /
"generate the pipeline report" (weekly-pipeline-report) vs "run the pipeline report"
(pipeline-intelligence-report); "update the pipeline" / "what does pipeline look
like" (weekly) vs "pipeline update" / "what's the pipeline look like" (intel).
Proposal: TRIM_DESC on weekly-pipeline-report — restrict its triggers to
"weekly pipeline report", "pipeline summary", "mid-month pipeline check",
"this week's numbers"; cede "pipeline report / pipeline update / pipeline look"
to pipeline-intelligence-report (which already owns the master-scoring scope).

1c. WARNING | TRIM_DESC | "deal-strategy-coach" vs "next-to-close"
deal-strategy-coach description: "asks which deals are likely to close";
next-to-close description: "which deals are most likely to close". Same ask,
two output shapes (coaching analysis vs ranked shortlist).
Proposal: TRIM_DESC on deal-strategy-coach — carve out close-ranking asks with an
explicit pointer to next-to-close (mirror of the existing next-to-close ↔
pipeline-intelligence-report disambiguation, which is correctly done in-body).

1d. INFO | REVIEW | "model-selection" vs "analysis-validator"
Not a phrase duplicate — a scope collision. Both claim unconditional ALWAYS:
model-selection "runs this skill at the start of every task, without exception …
before any planning, execution, or skill invocation"; analysis-validator "Always.
No exceptions." on every SignalForge quantitative output. Position ordering exists
only in body prose (model-selection = pre-plan; validator = post-analysis;
claim-compressor and signalforge-feedback sequence explicitly "after
analysis-validator and after signalforge-claim-compressor"). No manifest-level
precedence field exists.
Proposal: REVIEW — record explicit precedence (model-selection → analysis →
analysis-validator → signalforge-claim-compressor → signalforge-feedback) in one
place; do not rely on body prose alone.

================================================================
(2) CIRCULAR DELEGATION — 1 cycle
================================================================
CRITICAL | UPDATE_BODY | Cycle: "deal-strategy-coach" → "email-drafter" → "deal-strategy-coach"
  deal-strategy-coach, Manager-to-prospect email frameworks section:
    "When drafting manager-to-prospect emails, use the `email-drafter` skill"
  email-drafter, description (final sentence):
    "For deal strategy, diagnosis, or coaching (not email drafting), use
     deal-strategy-coach instead."
Mutual delegation = A→B→A. Extension of the same cycle: comms-drafter →
deal-strategy-coach ("this skill drafts, that skill diagnoses") → email-drafter →
back to deal-strategy-coach. No other cycle found; pipeline-intelligence-report →
closed-lost-analysis is one-way (closed-lost-analysis only names
pipeline-intelligence-report as an inbound caller, "called from
pipeline-intelligence-report", Mode 4 — not an outbound delegation).
Proposal: UPDATE_BODY on email-drafter — convert its back-reference into a
non-delegating lane note ("draft only; return to caller"), breaking the back-edge.

================================================================
(3) DANGLING DELEGATION TARGETS (named)
================================================================
CRITICAL | REVIEW | 15 named targets referenced in bodies, absent from the 14-file set
Execution-blocking (skills fail mid-run without these):
  signalforge-reports          — pipeline-intelligence-report Phase 5 + weekly-pipeline-report
                                 Step 4 mandate reading SKILL.md, DESIGN-SYSTEM.md,
                                 signalforge.css "BEFORE writing any HTML"
  bonusly-brand                — comms-drafter Step 0, email-drafter, sales-forecast,
                                 signalforge-claim-compressor ("use bonusly-brand instead")
  references/data-sources.md, references/report-structure.md, references/cadence.md,
  references/report-template.html, references/TEMPLATE_README.md
                                 — sales-forecast (5 mandatory reads)
  references/report-spec.md, references/queries.md — weekly-pipeline-report
  /mnt/skills/public/xlsx/scripts/recalc.py — stale-pipeline-report validation gate
Advisory (correctness of claims depends on them):
  prospect-research-multithreading — comms-drafter, email-drafter, deal-strategy-coach
  bonusly-data-questions, bonusly-product-questions, bonusly-business-reporting-questions,
  bonusly-rewards-questions, bonusly-ppp-questions, bonusly-feature-flag-questions,
  bonusly-deal-desk-questions, bonusly-datadog-questions — analysis-validator §12.4
  skill-orchestrator, CUSTOMER_DATA_REFERENCE, HUBSPOT_CONNECTOR_REFERENCE,
  SIGNALFORGE_PRODUCT_INSIGHT_SKILL — analysis-validator §11 Three-Way Sync;
                                     signalforge-feedback Activation Checklist
  caveman — signalforge-claim-compressor "Relationship to Caveman Skill"
Verified present in-set (not dangling): analysis-validator, closed-lost-analysis,
deal-strategy-coach, email-drafter, pipeline-intelligence-report, model-selection.
Note: these targets may exist outside the provided data; not verifiable here.
Proposal: REVIEW — build a resolution table mapping each named target to an
existing file or a registered stub, priority order: signalforge-reports →
bonusly-brand → sales-forecast/weekly-pipeline-report references/* → the rest.

================================================================
(4) VERSION CONFLICT
================================================================
WARNING | UPDATE_BODY | "analysis-validator" — v3.6 vs v3.2, survivor = v3.6
Evidence:
  - Body header: "Version: 3.6"; changelog top row: 3.6 | May 9, 2026
  - Body §7 Validation Trail template: "Validator: analysis-validator v3.2"
    (stale literal left from the 3.2-era template)
  - External cross-reference agrees with 3.6: pipeline-intelligence-report footer
    "✓ SignalForge Validated · Analysis Validator v3.6"
Count: 3.6 asserted 3 places vs v3.2 asserted 1 place → v3.6 survives.
Proposal: UPDATE_BODY on analysis-validator §7 trail template literal (v3.2 →
matching current version). No skill deletion warranted.
Not a conflict (recorded to avoid false positives): pipeline-intelligence-report
frontmatter "version: v6 · May 2026" vs "v4 Component Vocabulary" / v4 design
system — v4 is the signalforge-reports design-system version, a different axis.
Manifest has no version column, so cross-row version conflicts cannot be detected
from the manifest alone (see 7).

================================================================
(5) DESCRIPTIONS OVER 1,024 CHARACTERS
================================================================
INFO | TRIM_DESC | Count = 0 of 14
Arithmetic over the description_chars column:
  656, 897, 996, 792, 965, 676, 945, 1004, 1006, 962, 1006, 708, 762, 656
  max = 1,006 → 1,006 < 1,024 → 0 rows exceed; headroom = 18 chars.
Nearest to cap (1,006, 98.2% of limit): "pipeline-intelligence-report" and
"signalforge-claim-compressor"; third: "partner-digest" (1,004).
Column verified against bodies on 4 spot-checks (YAML folded-scalar normalization):
  analysis-validator 656/656, model-selection 676/676, next-to-close 945/945,
  partner-digest 1004/1004 — delta +0 each. Column is accurate as given.
Proposal: TRIM_DESC (preventive) on "pipeline-intelligence-report" and
"signalforge-claim-compressor" to open headroom under the 1,024 cap — not a
current violation.

================================================================
(6) HARDCODED PAGE IDs / DATES / PERSON NAMES IN BODIES
================================================================
CRITICAL | REVIEW | "analysis-validator"
  Person names + IDs: full GTM roster §12.3 — Alaina Loori 82535637, Shealagh
  Coughlin 119069206, Bryce Harmon 119337721, Hugo Lindqvist 77260721, Dana
  Mercer 83155923, Alex Franklin 84342457, Cole Ingram 83155924, Gavin Porter
  1520255671, Colleen Perry, Ellie Barton, Ashley Reyer, Megan Franz, Elena
  Sinclair, Youssef Elkhateeb, Amanda Czenkus, Ben Castelli, Amani Phipps
  210200121, John Thomas 78303262, Yasmin Wahid 89062643; escalation names
  "Manish or Amani" (G1-K, §10); example IDs 83155923 / 150582537 / 1520255671.
  Dates: April 26 2026, May 4 2026, May 9 2026, "stale as of March 28, 2023",
  "as of May 2026", "Confirmed current as of May 4, 2026".
  Proposal: REVIEW — extract roster/owner-ID table to a dated data reference;
  bodies hold a lookup pointer, not the roster.

CRITICAL | REVIEW | "partner-digest"
  Page/space IDs: cloudId 73fe98de-a4a3-4869-9f8a-bb1eeed4cf7f, spaceId
  1958248479, Partnerships Digest folder 2286616609, pages 2286321666,
  2265382925, 2236940297, 2237825028, 2239365136, 2238283777; Slack user
  U03QLMBL7AR. Names: Amani Phipps, Kelli, Jen Lee, Hani, Bryce, Sara.
  Dates: May 16 2026, May 19 2026, June 2 2026, 2026-05-17 changelog,
  "Q2/Q3 2026" page title. Proposal: REVIEW — move IDs to config block.

WARNING | REVIEW | "sales-forecast"
  Confluence: spaceId 2232811524, cloudId 73fe98de-…, parent page 2232582148.
  Names: Alaina (§2A), Elena (changelog v1.1). Dates: April 27 2026,
  "Q3 2026 … July 9, 2026" example. Proposal: REVIEW — same config extraction.

WARNING | REVIEW | "signalforge-feedback"
  Page IDs: 2295136266 (Feedback Log), parent 2234417154, Build Log 2247295002,
  spaceId 2232811524, cloudId 73fe98de-…. Names in examples: "Gavin Porter Rep
  Diagnostic", "Lowe's Conversation Analysis". Proposal: REVIEW — config block.

WARNING | REVIEW | "stale-pipeline-report"
  Slack channel ID C0561C1JCPJ (#revops-team); HubSpot owner ID 55483190
  (Bonusly Support); org 1973303 in link pattern; example dates 5/7, 5/15, 5/19;
  changelog 2026-06-10. Proposal: REVIEW — channel/owner IDs to config.

WARNING | UPDATE_BODY | "weekly-pipeline-report"
  Person name in skill title: "Weekly Pipeline Report — Ben Lavin · Demand
  Generation · Bonusly". Spreadsheet IDs 1CLZeOsElVDF_LF0ZG_t2nfwvhnZ6bpwqM_nX3WEYzcw
  and 1ENuaEcCuLjdKhMvp8FK3Ys1ek5Aw9ZuOZhsHJJFoB_k. Hardcoded window
  "Q2 (April 1 – June 30, 2026; total ≈ 64–65)" and static Q1 2026 actuals
  ($365,152 vs $475,000 plan = 77%; $2,490,532 vs $3,288,000 forecast = 76%) —
  violates the skill's own "always read live — do not hard-code values" rule and
  its quarter-agnostic posture. Proposal: UPDATE_BODY — replace the Q2-2026
  window with a computed quarter, and label the Q1 2026 actuals as a dated
  historical snapshot or remove.

WARNING | REVIEW | "deal-strategy-coach"
  Confluence page ID 2257879045 (AE Excellence Playbook, "April 2026" in URL);
  named routing persons "Perseus", "Farid" (ICP country/domain routing);
  dated pricing table "Pricing — 2026". Proposal: REVIEW — playbook URL + routing
  owners to a dated reference; pricing table flagged as date-bound data.

WARNING | REVIEW | "pipeline-intelligence-report"
  AE names + owner IDs (Bryce Harmon 119337721, Dana Mercer 83155923, Cole
  Ingram 83155924, Alex Franklin 84342457, Gavin Porter 1520255671 — "verified
  May 2026"); HubSpot org ID 1973303 in the required URL pattern; stage IDs
  150582536–1175632767. Data conflict inside the set: this roster lists 5 AEs
  while analysis-validator §12.3 declares "Core 6 AEs" including Hugo Lindqvist
  77260721 — the two hardcoded rosters disagree. Proposal: REVIEW — single
  roster source (see analysis-validator finding).

INFO | REVIEW | "closed-lost-analysis"
  Named example companies in taxonomy/interventions: Softheon, Estee Lauder,
  MinIO, LIFTOFF, Nestlé, Ozinga, Aurora Innovation, GCash, Ethos Cannabis,
  StickerYou; dated claims "30-deal AI-field sample from May 2026", "March 28,
  2023", "confirmed May 2026". Intentional illustrative examples — will rot.
  Proposal: REVIEW — label the block "illustrative, as of May 2026" or prune.

INFO | REVIEW | "model-selection"
  Registry date "last_checked: 2026-05-19" + model IDs — intentional and
  self-flagging (14-day staleness rule). Proposal: REVIEW — no change beyond
  honoring the existing staleness check.

INFO | REVIEW | "signalforge-claim-compressor"
  Example names Felix Construction, Panopto, Schneider Downs; changelog
  2026-05-09. Proposal: REVIEW — examples only; low rot risk.

INFO | REVIEW | "next-to-close"
  Org ID 1973303 in the deal URL pattern; stage IDs (shared constants).
  Proposal: REVIEW — same URL-pattern constant as pipeline-intelligence-report;
  fix once.

INFO | none | "comms-drafter", "email-drafter"
  No page IDs, no dates, no person names found in either body. Clean on this check.

================================================================
(7) MANIFEST DRIFT — BOTH DIRECTIONS
================================================================
INFO | REVIEW | No presence drift: 0 files with no manifest row, 0 manifest rows with no file
Arithmetic: 14 manifest rows vs 14 SKILL.md files provided; set difference in
both directions = 0. declared_name matches the frontmatter name and the filename
for all 14 (verified by inspection; e.g. pipeline-intelligence-report frontmatter
name = "pipeline-intelligence-report").
Secondary manifest gaps (schema, not presence):
  - No version column, though 7+ bodies carry versions (analysis-validator 3.6,
    pipeline-intelligence-report v6, partner-digest 1.1, sales-forecast 1.1,
    stale-pipeline-report 1.1, signalforge-claim-compressor 1.0) — the conflict
    in (4) is undetectable from the manifest alone.
  - No dependency column — findings (1)–(3) are undetectable from the manifest.
  - body_lines not re-counted (bodies supplied inline, not as files on disk).
Proposal: REVIEW — extend the manifest schema with version + depends_on columns
so findings (2)–(4) become manifest-detectable; presence drift requires no action.

================================================================
SUMMARY — one proposal per finding
================================================================
MERGE       comms-drafter + email-drafter (1a)
TRIM_DESC   weekly-pipeline-report triggers (1b); deal-strategy-coach triggers (1c);
            pipeline-intelligence-report + signalforge-claim-compressor desc headroom (5)
UPDATE_BODY email-drafter back-edge (2); analysis-validator v3.2 literal (4);
            weekly-pipeline-report hardcoded Q2-2026 window (6)
DELETE_SKILL none — no skill is dead or fully superseded on the provided data
REVIEW      precedence of ALWAYS gates (1d); dangling-target resolution table (3);
            config extraction for IDs/names (6, 8 skills); manifest schema (7)
Nothing rewritten — proposals only.
runbook-from-slack-threadPASS 1.00
Plain-English question: Turn the resolved incident thread into a runbook.
Prompt — what the model saw
Turn the resolved incident thread into a runbook: ordered steps with the exact command or action taken, who ran it, how success was verified, and rollback for each step that changed state. Every step must trace to a message in the thread. Add nothing the thread does not contain; where the thread is unclear, mark the step as needing confirmation.
cost $0.0052 · 51s · in 7,393 / out 2,101 tok
✓ ordered_steps✓ rollback✓ needs_confirmation✓ sourced

✓ Correct Answer (ground truth)

Plain-English answer: Ordered steps from the thread, each with the exact command, who ran it, how success was verified, and rollback where state changed.
The answer in one line: rule = no steps beyond what the thread contains; every step traces to a message id.
ambiguous_steporder=4; msg=M05; mark=needs-confirmation
ruleno steps beyond what the thread contains; every step traces to a message id.
stepsorder=1; msg=M01; actor=Bryce Harmon; action=acknowledge PagerDuty alert, take IC; state_change=False; verified_by=M01 | order=2; msg=M02; actor=Farid Osman; action=check queue depth; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M02 (48,213 pending) | order=3; msg=M04; actor=Farid Osman; action=pause auto-enqueue via feature flag; command=bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'; state_change=True; rollback=bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'; verified_by=M07 (queue depth falling) | order=4; msg=M05; actor=Elena Sinclair; action=cleared the dead set; state_change=True; needs_confirmation=thread gives no command or method for clearing the dead set | order=5; msg=M06; actor=Bryce Harmon; action=scale reward-worker to 6 replicas; command=kubectl scale deployment/reward-worker --replicas=6; state_change=True; rollback=kubectl scale deployment/reward-worker --replicas=3; verified_by=M07/M08 | order=6; msg=M08; actor=Cole Ingram; action=verify queue drained and error rate baseline; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M08 (returns 0) … (+2 more)
graded checksordered_steps · rollback · needs_confirmation · sourced

Model's actual answer

show the model's full answer
RUNBOOK — reward-worker queue backlog / Redis timeouts
Incident: #incident-2026-08-29-reward-queue (resolved)
Source: incident_thread.md (2026-08-29 export). Every step cites a thread message. Nothing added beyond the thread.

=====================================================================
STEP 1 — Acknowledge alert, take incident command
Trace: [M01] 2026-08-29 14:02:10Z
Who: Bryce Harmon
Action: PagerDuty alert fired for reward-worker queue depth > 10k; acknowledged and took IC. No command recorded.
Success verification: None documented — needs confirmation.
State change: Yes (incident ownership). Rollback: Not documented — needs confirmation. (No reversal or handoff of IC is recorded in the thread.)

STEP 2 — Measure queue depth
Trace: [M02] 2026-08-29 14:04:33Z
Who: Farid Osman
Command: `bundle exec rake sidekiq:queue_depth`
Result (reported): reward queue at 48,213 pending jobs. Stated normal is under 500.
Success verification: The command's own output is the measurement; no separate verification documented.
State change: No. Rollback: N/A.

STEP 3 — Inspect dead set
Trace: [M03] 2026-08-29 14:06:02Z
Who: Farid Osman
Action: Dead set has 112 jobs, all Redis::TimeoutError from around 13:58. Exact inspection command NOT recorded — needs confirmation.
Success verification: None documented — needs confirmation.
State change: No (reported as inspection only). Rollback: N/A.

STEP 4 — Pause enqueue (stop the bleed)
Trace: [M04] 2026-08-29 14:08:45Z
Who: Farid Osman
Command: `bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'`
Success verification: None directly documented. No flag-state readback is recorded. Later queue improvement ([M07], [M08]) is a reported observation and does NOT verify this flag change — needs confirmation.
State change: Yes.
Rollback (as documented in-thread for this step): `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`

STEP 5 — Clear the dead set
Trace: [M05] 2026-08-29 14:15:20Z
Who: Elena Sinclair
Action: Cleared out the dead set while in the console. Exact command NOT recorded — needs confirmation.
Success verification: None documented — needs confirmation.
State change: Yes (destructive — 112 dead jobs cleared per [M03] context).
Rollback: Not documented — needs confirmation. (No reversal is recorded for this step. This action is NOT established by the thread as an approved repeatable remediation.)

STEP 6 — Scale reward workers up
Trace: [M06] 2026-08-29 14:21:07Z
Who: Bryce Harmon
Command: `kubectl scale deployment/reward-worker --replicas=6` (was 3)
Success verification: None directly documented. No replica-count readback is recorded. [M07]/[M08] queue observations are later reported results, not isolated proof of this action's effect — needs confirmation.
State change: Yes.
Rollback (as documented in-thread for this step): `kubectl scale deployment/reward-worker --replicas=3`

STEP 7 — Monitoring observation (not an action)
Trace: [M07] 2026-08-29 14:33:41Z
Who: Farid Osman
Reported: Queue depth down to 9,400 and falling ~1,200/min. Measurement command NOT recorded — needs confirmation. (Values preserved as reported; no derived metrics computed.)

STEP 8 — Verify recovery
Trace: [M08] 2026-08-29 14:47:55Z
Who: Cole Ingram
Command: `bundle exec rake sidekiq:queue_depth`
Result (reported): returns 0. Also reported: error rate in Datadog back to baseline (Datadog query/command NOT recorded — needs confirmation).
Success verification: Queue-depth command output (0) plus reported Datadog baseline.
State change: No. Rollback: N/A.

STEP 9 — Re-enable enqueue
Trace: [M09] 2026-08-29 14:49:10Z
Who: Bryce Harmon
Command: `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`
Success verification: Documented — verified 40 new jobs processed cleanly in the next 3 minutes.
State change: Yes.
Rollback: Not documented — needs confirmation. (Do not infer the Step 4 disable command as this step's approved rollback.)

STEP 10 — Scale workers back down, resolve
Trace: [M10] 2026-08-29 14:55:00Z
Who: Bryce Harmon
Command: `kubectl scale deployment/reward-worker --replicas=3`
Success verification: Reported in the same message: queue stable at 0, incident resolved. No replica-count readback documented — needs confirmation.
State change: Yes.
Rollback: Not documented — needs confirmation. (Do not infer scaling back to 6 as this step's approved rollback.)

=====================================================================
ITEMS NEEDING CONFIRMATION (thread gaps)
- [M01] IC acknowledgement: no verification or reversal documented.
- [M03] Dead-set inspection: command absent.
- [M04] Flag disable: no direct flag-state verification.
- [M05] Dead-set clear: command, verification, and rollback all absent (destructive).
- [M06] Scale-up: no direct replica-count verification; effect not isolated in-thread.
- [M07] 9,400 / ~1,200/min: measurement command absent (numbers preserved as reported).
- [M08] Datadog baseline claim: query/command absent.
- [M09] Re-enable enqueue: no rollback documented.
- [M10] Scale-down: no rollback documented; "queue stable at 0" is a report, not a readback.
datadog-incident-timelinePASS 1.00
Plain-English question: Sequence the failure from the logs — what broke first, and what cascaded after it?
Prompt — what the model saw
From the log slice, sequence the failure: first error with timestamp, the cascade in order, the service and job involved, and the Datadog query that would confirm the first error. State what the logs do not show.
cost $0.0056 · 66s · in 5,096 / out 3,799 tok
✓ first_error_ts✓ service_named✓ cascade_ordered✓ query✓ gaps_stated

✓ Correct Answer (ground truth)

Plain-English answer: A reward-service Redis timeout at 14:01:12Z cascades to sidekiq at 14:01:40 and api-gateway at 14:03:05, with the Datadog query that confirms the first error.
The answer in one line: job_class_involved = RewardGiveJob
cascade_orderservice=reward-service; first_error_at=2026-09-03T14:01:12Z | service=sidekiq; first_error_at=2026-09-03T14:01:40Z | service=api-gateway; first_error_at=2026-09-03T14:03:05Z | service=web-app; first_error_at=2026-09-03T14:03:30Z
clean_servicespostgres
confirming_queryservice:reward-service status:error "@message:Redis::TimeoutError"
first_errortimestamp=2026-09-03T14:01:12Z; service=reward-service; message=Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s
job_class_involvedRewardGiveJob
logs_do_not_showroot cause of the Redis timeout, redis-server's own logs, infrastructure metrics (CPU/memory/network), any deploy or config change preceding the incident, who applied mitigation or when (no actor fields in logs)
graded checksfirst_error_ts · service_named · cascade_ordered · query · gaps_stated

Model's actual answer

show the model's full answer
SEQUENCE OF FAILURE — datadog_logs.csv + sidekiq_jobs.csv

FIRST ERROR
  2026-09-03T14:01:12Z | service "reward-service" | level ERROR
  "Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s"
  (earlier ERROR/WARN entries: none. Everything before 14:01:12 is INFO.)

CASCADE, IN ORDER (timestamps verbatim from the log slice)
  1. 14:01:12 reward-service ERROR — Redis connection to "redis-primary:6379" times out after 5s. [origin]
  2. 14:01:20 / 14:01:30 / 14:01:40 reward-service ERROR — "retry exhausted for RewardGiveJob" x3 (10s apart). Arithmetic: 14:01:12 -> 14:01:20 = 8s to first exhausted retry.
  3. 14:01:40 sidekiq ERROR — "RewardGiveJob failed: Redis::TimeoutError; retrying in 60s".
  4. 14:02:28 sidekiq ERROR — "RewardGiveJob failed: Redis::TimeoutError; retrying" (48s later).
  5. 14:02:30 sidekiq WARN — "Queue reward depth above 10,000". Arithmetic: 78s (1m18s) after first error.
  6. 14:03:05 api-gateway ERROR — "502 upstream timeout calling reward-service /gives" (first break on the user-facing path). Arithmetic: 113s (1m53s) after first error.
  7. 14:03:30 web-app ERROR — "Give form submission failed: upstream 502 from api-gateway". Arithmetic: 138s (2m18s) after first error.
  8. 14:03:31 -> 14:06:52 interleaved continuation:
       sidekiq RewardGiveJob failures: 14:03:31, 14:04:22, 14:05:26, 14:06:47 (4 more; 6 total incl. 14:01:40 and 14:02:28)
       api-gateway 502s: 14:03:48, 14:04:13, 14:05:16, 14:06:52 (4 more; 5 total incl. 14:03:05)
       web-app give failures: 14:04:45, 14:05:42, 14:06:49 (3 more; 4 total incl. 14:03:30)
     User-facing failure window in slice: 14:03:05 -> 14:06:52 = 227s (3m47s).
  9. 14:10:56 - 14:18:13 postgres INFO — "checkpoint complete" x6 (only entries in this window).
  10. 14:22:10 reward-service INFO — "Redis connection restored; resuming job processing". Arithmetic: 1258s (20m58s) after first error.
  11. 14:24:45 sidekiq INFO — "Queue reward depth below 500". Arithmetic: 155s (2m35s) after restore.

SERVICE / JOB
  Service: "reward-service" (dependency: "redis-primary:6379"; callers "api-gateway", "web-app"; worker "sidekiq").
  Job: "RewardGiveJob" on queue "reward" — named in every retry/failure line and in sidekiq_jobs.csv J-00001..J-00012 (12 of 16 failed jobs, 75%; 12/16 = 0.75).
  Also affected per sidekiq_jobs.csv only: "RecognitionDigestJob" J-00013..J-00016 (4 jobs), all "Redis::TimeoutError", failed_at 14:02:36 / 14:03:15 / 14:04:55 / 14:05:50.

DATADOG QUERY TO CONFIRM THE FIRST ERROR
  service:reward-service level:error "Redis::TimeoutError" "redis-primary:6379"
  with time range 2026-09-03T14:00:00Z -> 2026-09-03T14:02:00Z
  Expected hit: exactly one line at 14:01:12Z ("timed out after 5s"), before any "retry exhausted" line.

WHAT THE LOGS DO NOT SHOW
  - Root cause of the Redis timeout. No "redis-primary" logs, no CPU/memory/network/latency metrics. The 5s timeout is client-side; down vs partition vs saturation is undetermined.
  - Any correlation IDs: no trace/request IDs, host/pod/container names, tenant or user IDs.
  - Queue depth as data: only two threshold crossings ("> 10,000" at 14:02:30, "below 500" at 14:24:45). No measured peak, no point-in-time values, no backlog drain rate between them.
  - Coverage gaps: nothing between 13:59:30 and 14:01:12 (onset unknown), and no reward-service/api-gateway/web-app/sidekiq entries between 14:06:52 and 14:22:10 — whether user-facing errors continued during that ~15m window is not shown. Only postgres checkpoints appear there.
  - Linkage between the two files: no job_id in datadog_logs.csv, and the sidekiq_jobs.csv failed_at values (14:01:46-14:02:58 for "RewardGiveJob"; 14:02:36-14:05:50 for "RecognitionDigestJob") do not correspond to any timestamped log line. The 14:01:20/30/40 "retry exhausted" lines cannot be mapped to specific job IDs.
  - "RecognitionDigestJob" never appears in datadog_logs.csv — its failures are evidenced only by sidekiq_jobs.csv; no service/log line traces them.
  - Retry policy internals: the "retrying in 60s" at 14:01:40 vs actual next failure at 14:02:28 (48s gap) is not reconciled; no retry-count, backoff, or max-attempts fields.
  - Impact quantification: no counts of failed gives, affected users/accounts, or error rates — only individual log lines and 16 failed job rows.
  - What "resuming job processing" at 14:22:10 entailed (replay, drop, duplicate-give risk) and whether the 16 failed jobs were re-run successfully — no success/complete entries for J-00001..J-00016.
feature-flag-state-summaryPASS 1.00
Plain-English question: Summarize what each feature flag controls and who has it on.
Prompt — what the model saw
Summarize the state of the listed feature flags: what each controls per the code excerpt, which companies or segments have it on, targeting rules, and any flag with no code reference or no targets. Company counts must match the export. Do not describe a flag that is not in the export.
cost $0.0021 · 15s · in 3,473 / out 643 tok
✓ orphans✓ no_targets✓ counts_match✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: Recognition streaks on for 42 beta companies, redeem-flow redesign off with a 12-company target list, plus any flag with no code reference or no targets.
The answer in one line: rule = company counts must equal the export; do not describe flags absent from the export.
code_referencedrecognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, redeem_flow_redesign, analytics_dashboard_v3, ms_teams_app_v2
flags{"recognition_streaks_v2": {"state": "on", "targeting_rules": "segment:beta_companies", "company_count": 42}, "points_budget_guardrails": {"state": "on", "targeting_rules": "all_companies", "company_c
flags_with_no_targetspaused_offboard_cleanup
orphan_flags_no_code_referencelegacy_give_modal, survey_boosters_q3, paused_offboard_cleanup
rulecompany counts must equal the export; do not describe flags absent from the export.
graded checksorphans · no_targets · counts_match · no_invented

Model's actual answer

show the model's full answer
FEATURE FLAG SUMMARY (9 flags in export; 6 referenced in code excerpt)

FLAGS WITH CODE REFERENCES (6)

1. recognition_streaks_v2 — ON
   Controls: app/models/recognition.rb — records a give via StreakTracker when enabled.
   Targets: segment:beta_companies — 42 companies.

2. points_budget_guardrails — ON
   Controls: app/services/budget_service.rb — enforces point budget limits via BudgetService#enforce! on the giver's points.
   Targets: all_companies — 220 companies.

3. slack_dm_nudges — ON
   Controls: app/jobs/nudge_job.rb — gates the entire nudge job; when on, sends Slack DM nudges via SlackDm.send_nudge.
   Targets: segment:region_na — 87 companies.

4. redeem_flow_redesign — OFF
   Controls: app/controllers/redeem_controller.rb — renders RedeemV2Component when on, falls back to RedeemV1Component when off.
   Targets: targeted_list — 12 companies.

5. analytics_dashboard_v3 — ON
   Controls: app/controllers/analytics_controller.rb — loads AnalyticsV3 as the dashboard. No else-branch shown in excerpt; behavior when off not specified.
   Targets: segment:tier_three — 65 companies.

6. ms_teams_app_v2 — OFF
   Controls: app/services/teams_installer.rb — installs TeamsAppV2 when enabled.
   Targets: targeted_list — 9 companies.

FLAGS WITH NO CODE REFERENCE (3)

7. legacy_give_modal — OFF
   Targets: segment:legacy_plan — 14 companies. No code reference in the excerpt; what it controls is not determinable from provided data.

8. survey_boosters_q3 — ON
   Targets: segment:legacy_plan — 7 companies. No code reference; behavior not determinable.

9. paused_offboard_cleanup — OFF
   Targets: none (empty targeting_rules) — 0 companies. No code reference and no targets.

COMPANY COUNT CHECK (must sum to export total)
42 + 220 + 87 + 12 + 65 + 9 + 14 + 7 + 0 = 456 company-flag pairings across 9 flags (counts are per-flag and may overlap across flags; export does not provide distinct company totals).

NOTES / DATA GAPS
- Segment definitions (beta_companies, region_na, tier_three, legacy_plan) and the contents of "targeted_list" are not provided; I cannot enumerate which specific companies are in each.
- "targeted_list" is described as a targeting rule with no accompanying company list in the export.
- paused_offboard_cleanup is the only flag with zero targets and zero code reference — effectively inert per the data given.
nda-triagePASS 1.00
Plain-English question: Triage each NDA green, yellow, or red — and quote the clause that decides it.
Prompt — what the model saw
Triage each NDA as GREEN (standard approval), YELLOW (counsel review), or RED (full legal review). Cite the clause that drives each classification by quoting it. Catch embedded non-solicits, non-competes, missing carve-outs, and unusual term or governing law. One line of reasoning per NDA.
cost $0.0024 · 26s · in 3,520 / out 916 tok
✓ green✓ yellow✓ red✓ non_solicit_caught✓ clause_quoted✓ correct_labels

✓ Correct Answer (ground truth)

Plain-English answer: NDA-1 green (standard mutual terms), NDA-2 yellow for an embedded 18-month non-solicit, with the driving clause quoted for every classification.
The answer in one line: non_solicit_in = NDA-2
checksclassification_per_nda, clause_quotes_exist_in_document, non_solicit_caught, carveouts_checked
classifications{"NDA-1": {"classification": "GREEN", "driving_reason": "standard mutual NDA: 2-year term, standard carve-outs, Delaware law", "driving_clause_quote": "Carve-outs: information that (a) is or becomes p
non_solicit_inNDA-2
graded checksgreen · yellow · red · non_solicit_caught · clause_quoted · correct_labels

Model's actual answer

show the model's full answer
NDA-1 — GREEN
Clause: "Carve-outs: information that (a) is or becomes publicly available through no breach, (b) was known prior to disclosure, (c) is received from a third party without duty of confidence, (d) is independently developed, or (e) is disclosed by law or court order" (§3) — all five standard carve-outs present, mutual, 2+3 year term, Delaware law, no license/exclusivity (§5); nothing non-standard to flag.

NDA-2 — YELLOW
Clause: "During the term of this Agreement and for eighteen (18) months thereafter, neither party shall solicit for employment or hire any employee of the other party with whom it came into contact under this Agreement" (§4) — embedded employee non-solicit/no-poach extends restraints 18 months past the NDA's purpose (standard confidentiality terms otherwise: mutual, 3-year term, full carve-outs, Delaware law), so counsel should scope-limit or strike it before signature.

NDA-3 — RED
Clause: "For a period of three (3) years following the Effective Date, Recipient shall not, directly or indirectly, engage in or provide services to any business that competes with Discloser's business" (§2) — a 3-year non-compete with no geographic, industry, or role limitation, stacked on one-way confidentiality (§1), a long 5-year term, zero carve-outs (§4, so even public/prior/independently developed info is unprotected), and exclusive Irish jurisdiction (§3); this needs full legal review and likely renegotiation.