← back to all models ← RevenueBench home
aion-labs

aion-labs/aion-3.5

Mean 0.966 · 28/33 perfect tests · $3.06 total · 33/40 tests completed — remaining tests discarded from study (endpoint timeouts after dedicated retries) · modalities in: text · out: text · each card: the prompt → the correct answer (gold) → this model's actual answer

What do the modalities mean?

Modalities in means what you can send this model: text, images, files, video, or audio. Modalities out means what it can send back. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what this model can handle in real work.

data-crm 4 tests
owner-resolutionPASS 1.00
Plain-English question: Every open deal has a blank owner name — resolve each one to a real person using the owners table.
Prompt — what the model saw
The deal extract below has an owner name field that is blank on nearly every row and an owner id field that is populated. Using the owners table, resolve each open deal to an owner name. Report: number of deals resolved, deals whose owner id has no match in the owners table (list them), the owner ids that map to archived/deactivated owners (list their names), and the total pipeline amount per resolved owner. Do not guess a name for any unmatched id.
cost $0.1524 · 321s · in 12,288 / out 16,746 tok
✓ resolved_count✓ unmatched_none✓ deactivated_flagged✓ jake_amount✓ justin_amount✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: All 156 deals resolve to six named owners; jake_amount maps to Bryce Harmon ($1,054,144) and justin_amount to Alex Franklin ($624,310).
The answer in one line: jake_amount = 1,054,144; justin_amount = 624,310; resolved = 156
jake_amount1,054,144
justin_amount624,310
resolved156
deactivated
per_ownerBryce Harmon=1,054,144; Alex Franklin=624,310; Dana Mercer=341,195; Cole Ingram=288,161.43; Farid Osman=4,134; Elena Sinclair=2,100
graded checksresolved_count · unmatched_none · deactivated_flagged · jake_amount · justin_amount · no_invented

Model's actual answer

show the model's full answer
DEAL-TO-OWNER RESOLUTION — deals_open.csv × owners.csv

1) DEALS RESOLVED
156 of 156 deals resolved. Every owner_id in the deal extract matched a row in owners.csv. The 6 distinct owner_ids found in the deals file (119337721, 84342457, 83155923, 83155924, 716654662, 701163055) all exist in the owners table.

2) UNMATCHED OWNER IDS
None. No deal's owner_id is absent from owners.csv, so there is no deal left unnamed and no name guessed.

3) ARCHIVED/DEACTIVATED OWNERS
owners.csv contains two archived owners — Gavin Porter (owner_id 1520255671, archived=true) and Hugo Lindqvist (owner_id 77260721, archived=true). Neither owner_id appears on any deal row in deals_open.csv. Therefore zero open deals map to archived/deactivated owners.

4) TOTAL PIPELINE PER RESOLVED OWNER (sum of amount per owner_id)

Bryce Harmon (119337721) — 35 deals
24000+19656+13500+7000+2520 = 66,676
240000+99000+72000+70000+63600+45000+1 = 589,601
21000+23400+13680+5502+8160+1+11400+1 = 83,144
36000+31500+6000+10800+30275+17400+12600+18000+37440+18828+2880+36000+20880+10920+25200 = 314,723
66,676 + 589,601 + 83,144 + 314,723 = 1,054,144
Total: $1,054,144

Alex Franklin (84342457) — 67 deals
14850+13770+11200+9000+6360+5400+3240+2484+1920+1080 = 69,304
7200+19000+2880+1400+4800+1632+10000+9300+2700+2160 = 61,072
1800+3600+3840+15000+1968+4000+3600+4800+3120+2520 = 44,248
9000+2400+62000+5400+5100+16700+4400+1620+2600+7200 = 116,420
18000+17000+8316+8100+18000+12600+24000+15000+9000+7200 = 137,216
3780+16200+7200+4680+1800+18000+2730+2400+3060+18000 = 77,850
12000+1800+4400+31200+7200+1600+60000 = 118,200
69,304 + 61,072 + 44,248 + 116,420 + 137,216 + 77,850 + 118,200 = 624,310
Total: $624,310

Dana Mercer (83155923) — 24 deals
11250+10500+9000+9000+5400+4800+4600+1920 = 56,470
15000+4200+18900+27000+43875+20000+60000+8100 = 187,075
16250+3150+5000+2100+23400+5400+7350+25000 = 97,650
56,470 + 187,075 + 97,650 = 341,195
Total: $341,195

Cole Ingram (83155924) — 22 deals
2760+1330+7225.4+3360+5616+2700+7781.2+6947.5 = 37,720.10
40000+12168+2480.4+11116+11193+1875+31750 = 110,582.40
32175+4779.88+3334.8+58529.25+4140+18000+18900 = 139,858.93
37,720.10 + 110,582.40 + 139,858.93 = 288,161.43
Total: $288,161.43

Farid Osman (716654662) — 7 deals
1249+1+1+1+2880+1+1 = 4,134
Total: $4,134

Elena Sinclair (701163055) — 1 deal
2100
Total: $2,100

Cross-check: 35+67+24+22+7+1 = 156 deals; 1,054,144 + 624,310 + 341,195 + 288,161.43 + 4,134 + 2,100 = $2,314,044.43 total open pipeline.

Notes on data quality (stated, not inferred): the owners table contains 8 owners but only 6 appear on open deals; the two unused owners are the archived ones (Gavin Porter, Hugo Lindqvist). The final row of the CSV (Deal-57FF13) is missing the source field value (row has one fewer field); the deal's owner_id, amount, and stage are intact, so it is included in Elena Sinclair's totals.
gong-hubspot-join-integrity0.75
Plain-English question: How many call records point at deal ids that do not exist in the open pipeline?
Prompt — what the model saw
Check the integrity of the call-to-deal mapping below. The gong table lists calls in the last 90 days keyed to HubSpot deal ids; deals_open lists the current open default-pipeline deals. Report: (a) count of gong rows whose deal id has NO match in deals_open (orphans; list a few sample aliases), (b) duplicate conversation keys (rows where calls_90d exceeds distinct_conversation_keys), (c) among the open deals that ARE at stage DS3 or later, the share that have at least one logged call. Give sample aliases for the orphan finding, all taken from the data. Note: gong_owner_id is a Gong user id, not a HubSpot owner id.
cost $0.1852 · 383s · in 15,816 / out 20,439 tok
✓ orphan_count✗ dup_keys_0✓ ds3_coverage✓ sample_ids_present

✓ Correct Answer (ground truth)

Plain-English answer: 30 Gong call rows are orphans keyed to dead deal ids, none are duplicates, and DS3 call coverage is 29.4%.
The answer in one line: orphans = 30; ds3_cov_pct = 29.4; dup = 0
orphans30
ds3_cov_pct29.4
dup0
orphan_aliasesDeal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42, Deal-36EA09
graded checksorphan_count · dup_keys_0 · ds3_coverage · sample_ids_present

Model's actual answer

show the model's full answer
(a) ORPHAN GONG ROWS (deal id not found in deals_open)

30 of 67 gong rows are orphans. Arithmetic: 67 gong rows - 37 rows whose hs_deal_id matches a deals_open deal_id = 30 orphans (30/67 = 44.8% of the gong table).

Sample orphan aliases (exactly as given): Deal-8FA85D (46 calls), Deal-8FC3F9 (24), Deal-3B7945 (21), Deal-42B265 (21), Deal-9CCC42 (17), Deal-B038F0 (5), Deal-AC944F (5), Deal-38CA53 (5), Deal-422BA6 (6), Deal-C00480 (4). The other 20 orphans follow the same pattern (no matching deal_id in deals_open).

(b) DUPLICATE CONVERSATION KEYS

0 rows. In every one of the 67 gong rows, calls_90d equals distinct_conversation_keys (e.g., Deal-C6FE92: 11 = 11; Deal-8FA85D: 46 = 46; Deal-3B6668: 4 = 4). No row shows calls_90d > distinct_conversation_keys, so there is no evidence of duplicate conversation keys.

(c) DS3+ OPEN DEALS WITH AT LEAST ONE LOGGED CALL

Denominator: open deals at DS3 or later = 61 at DS3 + 14 at DS4 + 10 at DS5 = 85. (Check: 85 + 39 DS2 + 32 DS1 = 156 total open deals, matching the file.)

Numerator: of the 37 gong rows that matched an open deal, 25 are at DS3+ (12 DS3, 6 DS4, 7 DS5). The other 12 matched deals sit at DS1 or DS2 (e.g., Deal-03BA69 DS1, Deal-AF932D DS2), so they don't count here. 25 + 12 = 37, consistent with (a).

Share: 25 / 85 = 0.294, i.e. 29.4% of DS3+ open deals have at least one logged call in the last 90 days. The remaining 60 of 85 DS3+ deals (70.6%) have no call row in the gong table.

Notes and caveats:
- "At least one logged call" is measured by presence in the gong table; every matched row has calls_90d >= 3, so all 25 are unambiguous.
- Per your flag, gong_owner_id was not joined against deals_open owner_id; the match was done strictly on deal id.
- One integrity anomaly worth flagging beyond your three questions: several orphan rows carry last_call dates after the dataset's stated 90-day window and even future-dated calls (e.g., Deal-8FA85D last_call 2026-11-05 with 46 calls; Deal-42B265 last_call 2026-11-25; Deal-9CCC42 last_call 2026-11-27), and those same rows show calls_with_ds2_date = 0 and calls_with_sqo_date = 0 despite high call volumes. This suggests the orphan block may be closed/non-default-pipeline deals or a mis-keyed join rather than random noise, but the data provided does not say which.
snowflake-sql-trapped-catalogPASS 1.00
Plain-English question: Write one Snowflake query for first-month usage per customer using only the provided, trap-laden catalog.
Prompt — what the model saw
Using ONLY the catalog provided, write one Snowflake SQL query returning, per customer company, for its first calendar month as a customer: unique givers, recognition count, successful redemption count. Read the catalog notes carefully: some tables are stale or unpopulated and must be avoided, and one documented filter (excluding deleted givers) must NOT be applied to historical giving counts. Put a comment above each FROM or JOIN explaining the table choice. Return only fenced SQL.
cost $0.0791 · 225s · in 0 / out 10,676 tok
✓ avoids_stale_deals✓ avoids_unpopulated_live✓ uses_cohort_summary✓ succeeded_filter✓ no_giver_deletion_filter✓ returns_sql

✓ Correct Answer (ground truth)

Plain-English answer: A single correct SQL query that uses the right catalog tables and avoids the trap columns — graded on the exact tables and columns.
The answer in one line: use_table = PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
avoid_stalePRODUCTION.HUBSPOT.DEALS (stale)
avoid_unpopulatedV2_LIVE.OBJECTS_DEALS (unpopulated)
forbidden_filterdeleted-giver exclusion must NOT be applied
outputone fenced SELECT with per-FROM/JOIN table-choice comment
required_filtersucceeded redemptions only
use_tablePRODUCTION.PLG.COMPANY_COHORT_SUMMARY
graded checksavoids_stale_deals · avoids_unpopulated_live · uses_cohort_summary · succeeded_filter · no_giver_deletion_filter · returns_sql

Model's actual answer

show the model's full answer
```sql
-- ============================================================================
-- Per customer company, for its FIRST CALENDAR MONTH AS A CUSTOMER:
--   unique givers, recognition count, successful redemption count.
--
-- DATA GAPS — STATED EXPLICITLY (the catalog provides no alternative source):
--   1. NO raw recognition/giving events table exists in this catalog. The only
--      documented giving data are the precomputed M1_* columns on
--      PRODUCTION.PLG.COMPANY_COHORT_SUMMARY, so recognition count = M1_GIVES.
--   2. "Unique givers" is not a documented metric anywhere in this catalog.
--      M1_USERS (month-1 users) is the closest documented field and is used
--      as a PROXY; the catalog does not state that M1_USERS counts only users
--      who actually gave.
--   3. PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 is the
--      documented redemption source (STATE = 'succeeded' rows only), but the
--      catalog documents NO company or date columns on it, so it cannot be
--      aggregated per company per month without inventing schema. M1_REDEMPTIONS
--      is therefore the only documented per-company month-1 redemption count;
--      the catalog does NOT state whether it is already restricted to
--      STATE = 'succeeded' — verify before treating it as finance-grade.
--   4. The catalog does not document the company identifier column for
--      COMPANY_COHORT_SUMMARY; COMPANY_ID below is a PLACEHOLDER — confirm
--      the real column name before running.
--   5. The catalog does not document the M1 anchor (signup vs. first payment
--      vs. calendar month). The month label assumes M1_* corresponds to the
--      first calendar month as a customer — verify.
--
-- TABLES DELIBERATELY AVOIDED (per catalog notes):
--   - HUBSPOT_HUB_1973303.V2_LIVE.OBJECTS_DEALS   (unpopulated — do not use)
--   - PRODUCTION.HUBSPOT.DEALS                    (stale, last sync 2023-03)
--   - PRODUCTION.HUBSPOT.GONG_HUBSPOT_MAP_FAST    (retired — do not use)
--   - PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 (see gap 3)
--
-- DELETED-GIVER RULE: the documented filter excluding deleted givers must NOT
-- be applied to historical giving counts. No giver-level table exists in this
-- catalog, so no such filter is (or can be) applied here; M1_GIVES is consumed
-- as-is, consistent with that rule.
--
-- ARITHMETIC: none is required — all three metrics are precomputed counts in
-- the source table; this query performs no aggregation or derivation.
-- ============================================================================

SELECT
    -- PLACEHOLDER: company key column not documented in the catalog (gap 4).
    cs.COMPANY_ID                                                     AS company_id,
    -- First calendar month as a customer = calendar month containing the
    -- first subscription payment (the documented customer-start event).
    CAST(DATE_TRUNC('month', cs.FIRST_SUB_PAYMENT_DATE) AS DATE)      AS first_customer_month,
    -- PROXY for unique givers (gap 2): month-1 users, not verified givers.
    cs.M1_USERS                                                       AS unique_givers,
    -- Recognition count (gap 1): precomputed month-1 gives, with NO
    -- deleted-giver filter applied (per the catalog business rule).
    cs.M1_GIVES                                                       AS recognition_count,
    -- Successful redemption count (gap 3): precomputed month-1 redemptions;
    -- succeeded-only restriction not verifiable from the catalog.
    cs.M1_REDEMPTIONS                                                 AS successful_redemption_count
FROM
    -- PRODUCTION.PLG.COMPANY_COHORT_SUMMARY: one row per self-serve company
    -- with precomputed month-1 (M1_*) usage metrics. Chosen because it is the
    -- ONLY table in this catalog that documents per-company recognition and
    -- redemption counts; the remaining usable tables (HubSpot deals/owners/
    -- contacts/engagements, Gong map/transcripts, Chargebee subscriptions)
    -- contain sales and billing data, not giving or redemption events tied
    -- to a company's first customer month.
    PRODUCTION.PLG.COMPANY_COHORT_SUMMARY                             AS cs
WHERE
    -- A company is a customer only from its first subscription payment;
    -- companies without one have no "first calendar month as a customer."
    cs.FIRST_SUB_PAYMENT_DATE IS NOT NULL
;
```
crm-completeness-enrichmentPASS 1.00
Plain-English question: Audit the CRM for missing and invalid fields, and propose the fixes.
Prompt — what the model saw
Audit the CRM extract for completeness and propose fixes. Deals need owner, stage, amount, close date, why-buys; companies need industry, employee count, HQ country; contacts need email, title, persona. Report completeness percent per field. Identify duplicate company clusters (name variants, shared domain) with a survivor per cluster. List invalid emails and domain mismatches. Fill missing company fields only where the enrichment export has a matching row; where CRM and enrichment disagree list both and recommend a source. Never invent a value. End with the 10 fixes with the most pipeline amount at stake.
cost $0.1023 · 271s · in 1,642 / out 13,722 tok
✓ dup_clusters✓ invalid_emails✓ survivor_language✓ completeness_pct✓ no_invent

✓ Correct Answer (ground truth)

Plain-English answer: 9 companies missing employee count, 6 missing HQ country, 4 invalid emails, 1 domain mismatch, 2 duplicate clusters, and 25 enrichment matches.
The answer in one line: n_contacts = 52; n_companies = 34; enrichment_matches = 25; invalid_email_count = 4; n_dup_clusters = 2; domain_mismatch_count = 1 (+4 more below)
n_contacts52
domain_mismatch_count1
enrichment_matches25
invalid_email_count4
n_companies34
n_dup_clusters2
company_missingindustry=0; employee_count=9; hq_country=6
domain_mismatchesCT-0011
duplicate_clusters['C-0A092931', 'C-0A092932'], ['C-0A092933', 'C-0A092934']
invalid_emailsCT-0010, CT-0080, CT-0081, CT-0192
graded checksdup_clusters · invalid_emails · survivor_language · completeness_pct · no_invent

Model's actual answer

show the model's full answer
DATA QUALITY AUDIT — CRM EXTRACT
Files provided: companies.csv (33 rows), contacts.csv (52 rows), zoominfo_enrichment.csv (25 rows).
Deals file: NOT PROVIDED. Owner, stage, amount, close date, and why-buys cannot be audited — there are zero deal records. I will not invent any. This also means "pipeline amount at stake" cannot be computed; the final ranking uses company employee count as an explicit size proxy.

────────────────────────────────────────
1. COMPLETENESS BY FIELD
────────────────────────────────────────
COMPANIES (n = 33)
  industry:       33/33 = 100.0% populated (but 9 rows use non-standard values — see §6)
  employee_count: 24/33 = 72.7% (9 blank: C-EC3025, C-96039F, C-44EA29, C-D04904, C-B23205, C-60C75F, C-7BBDFA, C-50D386, C-93C8BF)
  hq_country:     26/33 = 78.8% (7 blank: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-93C8BF, C-EE9FFB)
  Post-enrichment fill (§5): employee_count 32/33 = 97.0%; hq_country unchanged at 78.8% (no fillable values exist).

CONTACTS (n = 52)
  email:  52/52 = 100% populated, but only 48/52 = 92.3% syntactically valid and 47/52 = 90.4% match their company domain (issues in §4)
  title:  38/52 = 73.1% (14 blank: CT-0000, CT-0022, CT-0072, CT-0080, CT-0081, CT-0092, CT-0120, CT-0121, CT-0122, CT-0132, CT-0141, CT-0162, CT-0170, CT-0192)
  persona: 36/52 = 69.2% (16 blank: CT-0000, CT-0022, CT-0041, CT-0060, CT-0070, CT-0080, CT-0081, CT-0082, CT-0092, CT-0110, CT-0132, CT-0162, CT-0171, CT-0172, CT-0180, CT-0181)

DEALS (n = 0 — file missing)
  owner, stage, amount, close date, why-buys: 0% — unauditable.

Coverage notes:
  Enrichment covers 25/33 companies = 75.8%. No ZI row for: ba969b.com, 332637.com, 93c8bf.com, ee9ffb.com, c9bb20.com, acme-corp.com (×2 records), globex.io (×2 records).
  Contacts exist for only 20 of 31 unique companies = 64.5%. No contacts at all for the two duplicate clusters or for C-2C60E5, C-B97B4E, C-2D7423, C-7BBDFA, C-50D386, C-BA969B, C-332637, C-93C8BF, C-EE9FFB, C-C9BB20.

────────────────────────────────────────
2. DUPLICATE COMPANY CLUSTERS
────────────────────────────────────────
Cluster 1 — shared domain acme-corp.com
  C-0A092931 (Technology, 500, US)
  C-0A092932 (tech, 510, USA)
  Survivor: C-0A092931 (earliest alias; both records equally complete; no ZI row to arbitrate). Merge C-0A092932 into it.
  Unresolved conflict: employee_count 500 vs 510 — enrichment has no acme-corp.com row, so no tiebreak exists. Flag for human verification; keep 500 until confirmed.

Cluster 2 — shared domain globex.io
  C-0A092933 (SaaS, 200, US)
  C-0A092934 (Technology, 200, US)
  Survivor: C-0A092933. Merge C-0A092934 into it.
  Unresolved conflict: industry SaaS vs Technology — no ZI row for globex.io; flag for human verification.

No other clusters: all remaining 29 domains are unique. Company aliases are opaque hashes, so domain is the only usable match key; no name-variant evidence exists. Note: C-7BBDFA and C-50D386 look similar (both "health care", Canada, blank headcount) but have different domains — not duplicates on the available evidence; do not merge.

Net: 33 records → 31 unique companies.

────────────────────────────────────────
3. INVALID EMAILS AND DOMAIN MISMATCHES
────────────────────────────────────────
Malformed (no domain after @) — 4:
  CT-0010  user0@   (C-66D1FC)
  CT-0080  user0@   (C-92D97D)
  CT-0081  user1@   (C-92D97D)
  CT-0192  user2@   (C-425E2A)
Domain mismatch — 1:
  CT-0011  user1@other-domain.com — contact's domain column and company C-66D1FC are both 66d1fc.com. "other-domain.com" matches no company in the extract. Verify whether this contact actually belongs to C-66D1FC.
Total: 5/52 = 9.6% of contacts have unusable emails. All other 47 emails match their company domain.

────────────────────────────────────────
5. FILLS FROM ENRICHMENT (CRM blank + matching ZI row with a value)
────────────────────────────────────────
All 8 fillable gaps are employee_count; every ZI row where CRM hq_country is blank is also blank (or has no ZI row), so zero country fills are possible.

  C-EC3025 (ec3025.com):   employee_count ← 400
  C-96039F (96039f.com):   employee_count ← 400
  C-44EA29 (44ea29.com):   employee_count ← 400  (hq_country still unfillable — ZI blank)
  C-D04904 (d04904.com):   employee_count ← 400  (hq_country still unfillable — ZI blank)
  C-B23205 (b23205.com):   employee_count ← 400
  C-60C75F (60c75f.com):   employee_count ← 400
  C-7BBDFA (7bbdfa.com):   employee_count ← 400
  C-50D386 (50d386.com):   employee_count ← 400

Unfillable (no value exists in either source — manual research required, do not fabricate):
  hq_country: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5 (ZI blank); C-93C8BF, C-EE9FFB (no ZI row)
  employee_count: C-93C8BF (no ZI row)

────────────────────────────────────────
6. CRM vs ENRICHMENT DISAGREEMENTS (both populated)
────────────────────────────────────────
Industry — CRM "tech"/"Technology"/"Tech " vs ZI "Computer Software" (10 domains):
  66d1fc.com (C-66D1FC), ec3025.com (C-EC3025), 44ea29.com (C-44EA29), 92d97d.com (C-92D97D), d04904.com (C-D04904), 77a95a.com (C-77A95A), aa8dda.com (C-AA8DDA), b25f40.com (C-B25F40), 60c75f.com (C-60C75F), 425e2a.com (C-425E2A)
  Recommendation: adopt ZoomInfo "Computer Software" as the canonical value — it is the standardized third-party taxonomy and the CRM values are inconsistent with each other anyway (3 rows "tech", 4 rows "Tech " with trailing space, 9 rows "Technology"). If "Technology" is an intentional broader bucket in your segment model, decide once and map both sources to it; either way the current mix is unusable for segmentation.

Employee count — no disagreements: where both sources are populated, all 25 overlap rows agree exactly (e.g., 900=900, 50=50, 1500=1500, 340=340). The only headcount conflict in the data is CRM-internal (acme 500 vs 510, §2).

Country — no factual disagreements, only formatting: US (9 rows) / USA (6 rows) / United States (2 rows) all denote the same country. Normalize to one value (recommend ISO "US"). ZI uses "United States" throughout.

Non-standard taxonomy inside CRM itself (no ZI conflict, but must be normalized): "health care" ×2 (C-7BBDFA, C-50D386 — ZI also says "health care", so ZI is no better here; normalize both to "Healthcare" per the 4 CRM rows already using it), "Tech " ×4, "tech" ×3, "SaaS" ×1 (C-0A092933).

────────────────────────────────────────
7. TOP 10 FIXES, RANKED BY STAKE
────────────────────────────────────────
Caveat repeated: no deal amounts exist in the provided data, so true "pipeline amount at stake" is uncomputable. Ranking below uses company employee count as the deal-size proxy plus breadth of impact. Re-rank the moment deals.csv is available.

 1. Obtain the deals extract. All five deal fields are 0% complete; the entire pipeline is unauditable and no amount-at-stake ranking is possible. Stake: 100% of pipeline (unquantified).
 2. C-EE9FFB (1,500 employees): hq_country blank, no ZI row — largest company with an unfillable gap; manual enrichment required.
 3. Merge acme-corp.com cluster (500/510 employees): survivor C-0A092931, absorb C-0A092932; resolve 500 vs 510 headcount by verification (no ZI tiebreak).
 4. C-66D1FC (900 employees): repair CT-0010 (malformed "user0@") and CT-0011 (domain mismatch) — two of three contacts at a large account are unreachable/unmatchable.
 5. Fill employee_count = 400 from ZI for C-EC3025, C-96039F, C-44EA29, C-D04904 (four 400-employee companies; grouped tie).
 6. Fill employee_count = 400 from ZI for C-B23205, C-60C75F, C-7BBDFA, C-50D386 (grouped tie).
 7. Merge globex.io cluster (200 employees): survivor C-0A092933, absorb C-0A092934; resolve SaaS vs Technology.
 8. C-93C8BF (size unknown): employee_count AND hq_country blank, no ZI row — the worst single record in the extract; manual enrichment required.
 9. Manually source hq_country for the five companies where both sources are blank: C-44EA29, C-D04904 (400 each post-fill), C-2C60E5 (340), C-2D1F1B (50), C-D73B89 (50).
10. Normalize taxonomy globally: industry (3 "tech" + 4 "Tech " + 9 "Technology" + 1 "SaaS" + 2 "health care" rows) and country (US/USA/United States) across all 33 records — silently corrupts every downstream segment and territory report.

Not in the top 10 because it cannot be size-weighted without deal data, but high volume: backfill 14 blank titles and 16 blank personas (30 blank contact cells, 57.7% of the blank-cell total), and pull enrichment for the 8 companies with no ZI row.
deal-intelligence 4 tests
pipeline-tieringPASS 1.00
Plain-English question: Score all 156 open deals into exactly one of LOCK / ACTION / BUILD / REVIVE / WATCH / RISKY.
Prompt — what the model saw
Score the open pipeline below on these signals: forecast category, stage, meetings_30d, contact count, and engagement recency. Assign exactly one tier per deal: LOCK, ACTION, BUILD, REVIVE, WATCH, or RISKY (RISKY = the forecast category disagrees with the engagement evidence). Never tier a deal LOCK if it has zero meetings_30d. Do NOT list every deal. Return ONLY a compact JSON object with exactly these keys: - "tier_counts": object mapping each tier to the number of deals in it (must sum to the total deal count), - "tier_examples": object mapping each tier to up to 3 example deal aliases from the data, - "risky_deals": array of the deal aliases you tiered RISKY, - "lock_violations": number of deals tiered LOCK that have zero meetings_30d (should be 0), - "pipeline_shape": one short paragraph on the overall pipeline shape. Note: inbound_emails_30d is 0 for every row due to a data defect, so treat meetings_30d as the inbound signal.
cost $0.2221 · 525s · in 17,462 / out 25,770 tok
✓ json_returned✓ tier_vocabulary✓ tier_counts_sum✓ risky_deals_listed✓ no_lock_violation_claimed

✓ Correct Answer (ground truth)

Plain-English answer: All 156 open deals tiered exactly once across the six tiers using forecast category, stage, meetings, contacts, and recency.
The answer in one line: total = 156
total156
graded checksjson_returned · tier_vocabulary · tier_counts_sum · risky_deals_listed · no_lock_violation_claimed

Model's actual answer

show the model's full answer
{
  "tier_counts": {"LOCK": 3, "ACTION": 26, "BUILD": 71, "REVIVE": 3, "WATCH": 47, "RISKY": 6},
  "tier_examples": {
    "LOCK": ["Deal-D348E1", "Deal-C26D20", "Deal-403845"],
    "ACTION": ["Deal-25F752", "Deal-C6FE92", "Deal-1FC049"],
    "BUILD": ["Deal-66D1FC", "Deal-93C8BF", "Deal-D73B89"],
    "REVIVE": ["Deal-2D1F1B", "Deal-333EBB", "Deal-57FF13"],
    "WATCH": ["Deal-E53952", "Deal-9AAE5F", "Deal-6787C2"],
    "RISKY": ["Deal-547B2B", "Deal-B7EBD1", "Deal-2465CE"]
  },
  "risky_deals": ["Deal-547B2B", "Deal-B7EBD1", "Deal-A2B47C", "Deal-2465CE", "Deal-584EE5", "Deal-FD9F4E"],
  "lock_violations": 0,
  "pipeline_shape": "156 deals total (3+26+71+3+47+6=156). The pipeline is heavily front- and mid-loaded: 71 BUILD (early-stage PIPELINE with real development motion, including large deals like Deal-66D1FC at $99,000 and Deal-EC3025 at $62,000) plus 47 WATCH (BEST_CASE/PIPELINE with email-only or stale engagement) make up 118 of 156 deals (~76%). Only 3 deals earn LOCK — all DS5 COMMIT with meetings_30d >= 1 and engagement within the last week — while 6 COMMIT deals are RISKY: zero meetings_30d despite a COMMIT forecast, four of them closing by 2026-09-11, which signals forecast inflation at the top of the funnel. 26 ACTION deals (BEST_CASE or active early-stage COMMIT with meetings/calls but not yet LOCK-qualified) are the near-term conversion candidates, and 3 REVIVE deals are past-due or dormant (Deal-2D1F1B, the $240,000 DS1 deal, last touched 2026-06-11; Deal-333EBB and Deal-57FF13 have close dates already elapsed). Note: inbound_emails_30d is 0 on every row per the stated data defect, so meetings_30d was used as the inbound signal throughout."
}
call-transcript-extractionPASS 1.00
Plain-English question: Extract the CRM write-back fields from each call transcript as JSON.
Prompt — what the model saw
For each transcript, extract CRM write-back fields as JSON: why-buys (prospect statements only), pain points, stakeholders from the speaker list, budget signal (prospect-stated or null), timeline signal, competitor mentioned (only if the prospect raised it), next step (only if explicitly agreed), objections, confidence. Never fill a field from what the rep said or hoped.
cost $0.1113 · 307s · in 846 / out 15,616 tok
✓ json_emitted✓ budget_signal✓ null_discipline✓ stakeholders✓ competitor_prospect_only

✓ Correct Answer (ground truth)

Plain-English answer: Per transcript: prospect-only why-buys, pain points, stakeholders, budget and timeline signals, and competitors — e.g. TX-001: '$40k earmarked', live before open enrollment.
The answer in one line: rep = Alex Franklin
allowed_newTX-001, TX-002, TX-003, TX-004, TX-005, TX-006, Achievers, Bucketlist, Nectar
per_transcript{"TX-001": {"why_buys": ["automating anniversary and birthday awards"], "pain_points": ["HR team of three cannot keep up with it manually", "track everything in a spreadsheet, and people slip through
repAlex Franklin
rep_sourced_nulls{"TX-003": ["budget_signal \u2014 only the rep mentioned pricing/budget"], "TX-004": ["next_step \u2014 rep proposed, prospect said 'no promises'"], "TX-006": ["budget_signal \u2014 rep decoy, prospec
graded checksjson_emitted · budget_signal · null_discipline · stakeholders · competitor_prospect_only

Model's actual answer

show the model's full answer
null = no prospect-stated signal in the transcript (nothing taken from the rep's statements).

```json
[
  {
    "transcript_id": "TX-001",
    "deal_alias": "Deal-CFE7F4",
    "why_buys": [
      "\"The big win for us would be automating anniversary and birthday awards\" (Prospect, VP People)"
    ],
    "pain_points": [
      "HR team of three cannot keep up with manual anniversary/birthday awards (Prospect, VP People)",
      "Recognition tracked in a spreadsheet; \"people slip through the cracks\" (Prospect, HR Admin)"
    ],
    "stakeholders": [
      "Prospect (VP People)",
      "Prospect (HR Admin)"
    ],
    "budget_signal": "$40k earmarked for engagement tools this fiscal year (Prospect, VP People)",
    "timeline_signal": "Live before open enrollment in November (Prospect, VP People)",
    "competitor_mentioned": "Achievers — evaluated last year; \"too heavy for a team our size\" (Prospect, VP People)",
    "next_step": "Security review with IT lead on September 12 — explicitly agreed (Prospect, VP People)",
    "objections": [
      "Needs SSO and audit logs for IT to sign off (Prospect, HR Admin)"
    ],
    "confidence": "high — budget ✓, timeline ✓, agreed next step ✓; the only named competitor (Achievers) was already rejected by the prospect"
  },
  {
    "transcript_id": "TX-002",
    "deal_alias": "Deal-70BB30",
    "why_buys": [
      "Tie recognition to retention for the hourly workforce (Prospect, Head of Total Rewards)"
    ],
    "pain_points": [
      "Regretted turnover over 30% in the hourly workforce (Prospect, Head of Total Rewards)"
    ],
    "stakeholders": [
      "Prospect (Head of Total Rewards)",
      "Prospect (CFO)"
    ],
    "budget_signal": "$25k pilot budget approved for this quarter (Prospect, CFO)",
    "timeline_signal": "Decision by end of September (Prospect, CFO)",
    "competitor_mentioned": null,
    "next_step": "Send pilot agreement; prospect will route it to legal this week — explicitly agreed (Prospect, CFO)",
    "objections": [
      "Workday integration \"has to be rock solid\" — stated as the CFO's one condition (Prospect, CFO)"
    ],
    "confidence": "high — budget ✓, timeline ✓, agreed next step ✓, CFO in the room; prospect stated this is the first vendor they've had a real demo with (no active competitor)"
  },
  {
    "transcript_id": "TX-003",
    "deal_alias": "Deal-530B50",
    "why_buys": [
      "Make recognition visible across 12 retail locations (Prospect, People Ops Manager)"
    ],
    "pain_points": [
      "Store managers have zero budget autonomy for on-the-spot recognition today (Prospect, People Ops Manager)"
    ],
    "stakeholders": [
      "Prospect (People Ops Manager)"
    ],
    "budget_signal": null,
    "timeline_signal": "\"Honestly there's no rush on our side until Q1\" (Prospect, People Ops Manager)",
    "competitor_mentioned": "Bucketlist — CEO used it at her last company and liked it (Prospect, People Ops Manager)",
    "next_step": "Schedule a call with the CEO; People Ops Manager will send two time options — explicitly agreed (Prospect, People Ops Manager)",
    "objections": [
      "\"The CEO has to be sold first — she decides anything people-related\" (Prospect, People Ops Manager)",
      "No urgency until Q1 (Prospect, People Ops Manager)"
    ],
    "confidence": "medium — agreed next step ✓, timeline ✓ (but deferred to Q1); budget ✗ (none stated); decision-maker (CEO) not yet engaged and has positive prior experience with a competitor (Bucketlist)"
  },
  {
    "transcript_id": "TX-004",
    "deal_alias": "Deal-180D02",
    "why_buys": [
      "Consolidate three separate recognition tools into one (Prospect, VP People)"
    ],
    "pain_points": [
      "Paying for three tools and none of them talk to their HRIS (Prospect, VP People)"
    ],
    "stakeholders": [
      "Prospect (VP People)",
      "Prospect (IT Security Lead)"
    ],
    "budget_signal": "Under $15k annually can be approved without going to the board (Prospect, VP People)",
    "timeline_signal": "Procurement cycle runs six to eight weeks minimum (Prospect, IT Security Lead)",
    "competitor_mentioned": null,
    "next_step": null,
    "objections": [
      "\"The security review took three months for our last vendor — that's my hesitation\" (Prospect, IT Security Lead)",
      "Procurement cycle runs 6–8 weeks minimum (Prospect, IT Security Lead)",
      "Proposed CFO follow-up not committed: \"Maybe — I need to check her calendar, no promises\" (Prospect, VP People)"
    ],
    "confidence": "low — budget threshold ✓, timeline ✓ (but 6–8 week procurement minimum), agreed next step ✗ (prospect explicitly declined to commit); explicit security-review hesitation on record. Positives: clear consolidation pain and a defined <$15k approval path"
  },
  {
    "transcript_id": "TX-005",
    "deal_alias": "Deal-F8767A",
    "why_buys": [
      "Automate service milestones (Prospect, HR Director)",
      "Analytics on recognition equity across departments (Prospect, HR Director)"
    ],
    "pain_points": [
      "Night-shift teams feel invisible (Prospect, People Ops Coordinator)",
      "Night-shift engagement scores run 20 points lower (Prospect, People Ops Coordinator)"
    ],
    "stakeholders": [
      "Prospect (HR Director)",
      "Prospect (People Ops Coordinator)"
    ],
    "budget_signal": "$12k approved under engagement line (Prospect, HR Director)",
    "timeline_signal": "Running before the January all-hands (Prospect, HR Director)",
    "competitor_mentioned": "Nectar — currently mid-pilot; \"you'd need to beat that experience\" (Prospect, HR Director)",
    "next_step": "Present to the exec team on October 2 — explicitly agreed (Prospect, HR Director)",
    "objections": [
      "Exec team skeptical after a failed rollout two years ago (Prospect, HR Director)",
      "Active Nectar pilot in place that this would need to beat (Prospect, HR Director)"
    ],
    "confidence": "medium-high — budget ✓, timeline ✓, agreed next step ✓ (dated exec presentation); offset by an active competitor pilot (Nectar) and stated exec-team skepticism"
  },
  {
    "transcript_id": "TX-006",
    "deal_alias": "Deal-EE195F",
    "why_buys": [
      "\"I'd love to cut the admin time on service awards\" (Prospect, HR Manager)"
    ],
    "pain_points": [
      "Personally spends five hours a month ordering and shipping plaques (Prospect, HR Manager)"
    ],
    "stakeholders": [
      "Prospect (HR Manager)"
    ],
    "budget_signal": "\"Budget isn't the issue — time is\" (Prospect, HR Manager) — no dollar amount stated",
    "timeline_signal": "Fiscal year resets in February; \"a Q1 start is realistic\" (Prospect, HR Manager)",
    "competitor_mentioned": null,
    "next_step": "Send one-page overview; HR Manager will forward it to the COO this week — explicitly agreed (Prospect, HR Manager)",
    "objections": [
      "\"Our COO usually prefers building things in-house\" (Prospect, HR Manager)",
      "Alternative under evaluation is an internal build, not another vendor: \"Nobody else — we're comparing this against just doing it internally\" (Prospect, HR Manager)"
    ],
    "confidence": "medium — budget ✓ (not a constraint), timeline ✓ (Q1), agreed next step ✓, no competing vendors; offset by build-vs-buy risk: the COO (not yet engaged) prefers in-house"
  }
]
```

Notes on exclusions per the rules: TX-003 budget is null because the only pricing figure ($8/employee/month) came from the rep. TX-004 next_step is null because the only proposed follow-up was answered "Maybe — no promises." TX-002 and TX-006 competitor fields are null — no vendor competitor was raised (TX-002: "first vendor we've had a real demo with"; TX-006: internal build, captured under objections). Stakeholders are prospect speakers only, per the speaker list; referenced non-speakers (IT lead in TX-001, legal in TX-002, CEO in TX-003, CFO in TX-004, exec team in TX-005, COO in TX-006) appear only where a prospect statement put them on record.
next-to-close0.75
Plain-English question: Which 3 deals are closest to signature right now?
Prompt — what the model saw
Which 3 deals are closest to signature right now? Use the CRM data and the Slack thread; Slack often has the real status before the CRM does. For each: deal, amount, why it is close, what is left. Under 250 words.
cost $0.0256 · 40s · in 181 / out 1,671 tok
✗ exactly_three✓ top3_named✓ slack_first✓ blocked_excluded

✓ Correct Answer (ground truth)

Plain-English answer: Deal-547B2B ($11,200), Deal-B7EBD1 ($9,000), and Deal-403845 ($9,000) — all DS5 COMMIT deals closing September 10-11, with what is left on each.
blocked_excludeddeal_alias=Deal-2465CE; reason=champion left + procurement freeze (Slack #sales-sparring 2026-09-04)
exactly_three1
ranking_ruleSlack signature-imminent deal ranks first; remaining ranked by stage (DS5>DS4), forecast (COMMIT>BEST_CASE>PIPELINE), close_date asc, amount desc
top3rank=1; deal_alias=Deal-547B2B; amount=11200; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11 | rank=2; deal_alias=Deal-B7EBD1; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-10 | rank=3; deal_alias=Deal-403845; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11
graded checksexactly_three · top3_named · slack_first · blocked_excluded

Model's actual answer

show the model's full answer
**3 deals closest to signature** (CRM + #sales-sparring thread, 2026-09-04):

**1. Deal-547B2B — $11,200** (Alex Franklin, DS5/COMMIT, close 2026-09-11)
- Why: Slack says redlines came back clean, the signing page is out, and their VP People said they are signing tomorrow — signature-imminent.
- Left: the signature itself.

**2. Deal-403845 — $9,000** (DS5/COMMIT, close 2026-09-11)
- Why: Slack says the order form is with their finance team — one step from signature.
- Left: finance sign-off on the order form.
- Data note: Slack has Dana Mercer calling it "my side," but CRM lists Alex Franklin as owner — discrepancy flagged as given; no resolution in the data.

**3. Deal-A2B47C — $6,360** (Alex Franklin, DS5/COMMIT, close 2026-09-11)
- Why: Slack confirms it's warm, at normal legal-review pace, with a near-term CRM close date. Deal-D348E1 ($13,770) sits at the same legal-review stage but closes 2026-10-15, so A2B47C ranks ahead.
- Left: legal review completion, then signature.

**Excluded:**
- Deal-2465CE ($5,400, COMMIT): Slack says champion left, procurement froze new vendors, pulled from commit — blocked, realistically Q4.
- Deal-B7EBD1 ($9,000, COMMIT, close 2026-09-10): no Slack mention; CRM-only, status unverified.

Combined top-3 value: $11,200 + $9,000 + $6,360 = **$26,560**.
product-gap-detectionPASS 1.00
Plain-English question: Which prospect complaints are true product gaps versus rollout or enablement issues?
Prompt — what the model saw
Surface true product gaps raised by prospects in active deals. For each candidate: quote the prospect line, classify as real gap, plan gate, or rollout/enablement issue using the product docs excerpt, and name the deal and amount. Only real gaps go in the summary. Rep language does not count as prospect voice.
cost $0.0243 · 34s · in 0 / out 1,575 tok
✓ real_gaps✓ plan_gate_classified✓ enablement_classified✓ gap_summary

✓ Correct Answer (ground truth)

Plain-English answer: Two real gaps (SCIM provisioning, ADP Workforce Now integration) separated from plan gates and enablement issues, each with the prospect quote and deal.
allowed_newTG-001, TG-002, TG-003, TG-004, TG-005, SCIM user provisioning, ADP Workforce Now integration, Custom report builder
enablement_issuestranscript_id=TG-004; deal_alias=Deal-84DBA6; amount=16000; feature=Slack integration recognition adoption; quote=We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.; why_enablement=Slack integration exists on all plans; managers were never trained
excluded_rep_voicetranscript_id=TG-005; line=Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap.; reason=rep voice, not prospect-raised
plan_gatestranscript_id=TG-003; deal_alias=Deal-CFE7F4; amount=32000; feature=Custom report builder; tier_required=Enterprise; quote=I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?
real_gapstranscript_id=TG-001; deal_alias=Deal-EC3025; amount=62000; feature=SCIM user provisioning; quote=We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.; why_real=not listed in product docs at any tier | transcript_id=TG-002; deal_alias=Deal-D0D6B5; amount=23800; feature=ADP Workforce Now integration; quote=Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.; why_real=supported HRIS list is Workday/BambooHR/Gusto/Rippling only
graded checksreal_gaps · plan_gate_classified · enablement_classified · gap_summary

Model's actual answer

show the model's full answer
CANDIDATE-BY-CANDIDATE ANALYSIS (5 prospect-facing lines found; 1 excluded as rep voice)

TG-001 — Deal-EC3025
Prospect line: "We need SCIM user provisioning — our security team will not approve any tool that requires manual account management."
Docs check: "SCIM user provisioning and ADP Workforce Now integrations are NOT currently listed as supported capabilities."
Classification: REAL GAP. Capability absent from all tiers (Core/Pro/Enterprise) and from the supported integrations list.
Amount: NOT PROVIDED — no amount field exists in the supplied data.

TG-002 — Deal-D0D6B5
Prospect line: "Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us."
Docs check: HRIS integrations listed are Workday, BambooHR, Gusto, Rippling (Pro and above); ADP Workforce Now explicitly NOT supported.
Classification: REAL GAP. Prospect states it is a dealbreaker; capability absent.
Amount: NOT PROVIDED.

TG-003 — Deal-CFE7F4
Prospect line: "I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?"
Docs check: "Custom report builder" = Enterprise tier only ("—" on Core and Pro).
Classification: PLAN GATE. Capability exists in-product on Enterprise; blocked by tier, not by absence. Excluded from summary.
Amount: NOT PROVIDED.

TG-004 — Deal-84DBA6
Prospect line: "We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it."
Docs check: Slack integration = available on all plans.
Classification: ROLLOUT/ENABLEMENT ISSUE. Product is live and functioning; the failure is training/adoption, not capability. Excluded from summary.
Amount: NOT PROVIDED.

TG-005 — Deal-36C33F
The mobile-app gap is raised by Alex Franklin (rep), not the prospect: "we don't have a native mobile app for hourly workers yet."
Prospect's own line: "The web version should be fine for our office staff for now." — acceptance, not a gap.
Classification: EXCLUDED — rep language does not count as prospect voice, and the prospect raised no gap.
Amount: NOT PROVIDED.

ARITHMETIC
Candidates reviewed: 5 prospect-facing exchanges (TG-001 through TG-005).
5 = 2 real gaps (TG-001, TG-002) + 1 plan gate (TG-003) + 1 enablement issue (TG-004) + 1 rep-voice exclusion (TG-005).

SUMMARY — TRUE PRODUCT GAPS (prospect-raised, absent from product docs)

1. SCIM user provisioning — Deal-EC3025. Prospect (IT Security Lead): "We need SCIM user provisioning — our security team will not approve any tool that requires manual account management." Not supported on any tier per docs snapshot. Amount: not provided in data.

2. ADP Workforce Now integration — Deal-D0D6B5. Prospect (HRIS Manager): "Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us." Not supported per docs snapshot. Amount: not provided in data.

DATA GAPS (explicit): No deal amounts, stages, or close dates were included in the provided files, so amounts cannot be reported for any deal. Docs are a snapshot dated 2026-08; roadmap commitments (e.g., Alex Franklin's "let me confirm where ADP stands on our roadmap") cannot be verified from the data provided.
rep-performance 2 tests
multithreading-gapPASS 1.00
Plain-English question: Which open deals are single-threaded or under-threaded?
Prompt — what the model saw
Find every open deal that is single-threaded (fewer than 2 active contacts) or under-threaded (fewer than 3, or all contacts in one persona). Active means engaged in the last 60 days and not marked former. Personas: economic buyer, champion, HR admin, IT security, finance. For each flagged deal: amount, stage, active contact count, personas present, personas missing, the most valuable persona to add given the stage, and any on-file unengaged contact at that company who fits; otherwise say none on file.
cost $0.1990 · 481s · in 5,472 / out 25,271 tok
✓ flagged_recall✓ persona_language✓ single_thread_term✓ on_file_fit

✓ Correct Answer (ground truth)

Plain-English answer: Deals with fewer than 2 active contacts (single-threaded) or fewer than 3, or all in one persona (under-threaded), each flagged with the missing persona.
The answer in one line: active_window_start = 2026-07-07
active_window_start2026-07-07
allowed_new59 items — e.g. C-077A0E, C-0D15DF, C-1BDFE0, C-2AE3AA
best_persona_ruleDS2/DS3: economic buyer > IT security > finance > HR admin > champion; DS4/DS5: economic buyer > finance > IT security > HR admin > champion; first persona not present among active contacts
flag_ruleflag if active_contacts < 2 (single-threaded) or active_contacts < 3 or all active contacts share one persona (under-threaded); active = engaged within 60 days and not former
flaggeddeal_alias=Deal-36C33F; amount=15000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-50D386; amount=36000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5408B0; amount=14850; stage=DS4; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5BFE3B; amount=18000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-885F45; amount=9300; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=IT security | deal_alias=Deal-92D97D; amount=60000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer … (+5 more)
healthy_not_flaggedDeal-4B0BEB, Deal-84DBA6, Deal-D348E1
persona_universeeconomic buyer, champion, HR admin, IT security, finance
graded checksflagged_recall · persona_language · single_thread_term · on_file_fit

Model's actual answer

show the model's full answer
DATA LIMITATIONS (stated up front): the provided files contain no deal amount, no stage, and no open/closed status field. I cannot report amount or stage for any deal, and "open" cannot be verified — all 14 deals in deal_contacts.csv were assessed. Because stage is missing, the "most valuable persona to add" below is a gap-based recommendation, not stage-informed.

METHOD
Active = last_engaged_date on or after 2026-07-26 (2026-09-24 minus 60 days) AND is_former = false. Everything else (former contacts, contacts lapsed before 2026-07-26) is excluded from active counts.

SINGLE-THREADED (fewer than 2 active) — 5 deals

1. Deal-EC3025 (61032318100, C-FDD0C7) — amount/stage: not provided
   Active: 1 of 2 (CT-047C54, champion, 2026-09-02). CT-F2C1AE (economic buyer) excluded: is_former=true.
   Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: CT-6827DB (Chief People Officer, economic buyer).

2. Deal-92D97D (59728118877, C-E23238) — amount/stage: not provided
   Active: 1 of 2 (CT-01F5B4, HR admin, 2026-08-28). CT-A902AE (champion) lapsed: 2026-06-01 < 2026-07-26.
   Personas present: HR admin. Missing: economic buyer, champion, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: none on file (no C-E23238 rows in unengaged_contacts.csv). Note: lapsed on-deal champion CT-A902AE is a re-engagement candidate.

3. Deal-36C33F (63739413805, C-077A0E) — amount/stage: not provided
   Active: 1 of 3 (CT-4FE556, IT security, 2026-08-15). CT-40B45 (champion) and CT-86B22F (economic buyer) excluded: is_former=true.
   Personas present: IT security. Missing: economic buyer, champion, HR admin, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: CT-1DB73E (Chief People Officer, economic buyer).

4. Deal-FCBE5B (62639586615, C-737030) — amount/stage: not provided
   Active: 1 of 1 (CT-4A5317, champion, 2026-08-29).
   Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: none on file (no C-737030 rows).

5. Deal-F9A08A (49757401138, C-0D15DF) — amount/stage: not provided
   Active: 1 of 2 (CT-931B10, champion, 2026-09-03). CT-913581 (economic buyer) lapsed: 2026-06-20 < 2026-07-26.
   Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: CT-697541 (Chief People Officer, economic buyer). Note: lapsed on-deal economic buyer CT-913581 is a re-engagement candidate.

UNDER-THREADED (fewer than 3 active, or all active in one persona) — 6 deals

6. Deal-50D386 (61055128146, C-EB10E4) — amount/stage: not provided
   Active: 2 of 2 (CT-AA41B2 champion 2026-09-01; CT-B9C35B HR admin 2026-08-25). 2 < 3.
   Personas present: champion, HR admin. Missing: economic buyer, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: CT-A1C4B3 (Chief People Officer, economic buyer).

7. Deal-D0D6B5 (60081655042, C-32918E) — amount/stage: not provided
   Active: 3 of 3 (CT-87CED4, CT-DE6D7C, CT-FD70B2 — all champions, 2026-09-02 / 2026-08-19 / 2026-08-07). Count is 3, but all in one persona.
   Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: CT-1FA4DB (Chief People Officer, economic buyer).

8. Deal-5BFE3B (51674270311, C-535D36) — amount/stage: not provided
   Active: 2 of 2 (CT-57123B, CT-5CE757 — both champions, 2026-08-31 / 2026-08-12). Triggers both rules: 2 < 3 AND single persona.
   Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
   Most valuable add: economic buyer.
   On-file unengaged fit: none on file (no C-535D36 rows).

9. Deal-885F45 (60686135564, C-5E8EFB) — amount/stage: not provided
   Active: 2 of 2 (CT-51C81E economic buyer 2026-08-26; CT-D9A0E8 champion 2026-08-11). 2 < 3.
   Personas present: economic buyer, champion. Missing: HR admin, IT security, finance.
   Most valuable add: IT security (economic buyer already engaged; among missing personas, security review is the typical gate).
   On-file unengaged fit: CT-B3F25D (IT Security Lead, IT security).

10. Deal-5408B0 (60182332309, C-2AE3AA) — amount/stage: not provided
    Active: 2 of 2 (CT-D33AE4 champion 2026-09-01; CT-8742FD HR admin 2026-08-18). 2 < 3.
    Personas present: champion, HR admin. Missing: economic buyer, IT security, finance.
    Most valuable add: economic buyer.
    On-file unengaged fit: CT-07FA76 (Chief People Officer, economic buyer).

11. Deal-C6D97A (62121783047, C-5A8FC2) — amount/stage: not provided
    Active: 3 of 3 (CT-223DDC, CT-B03555, CT-4E8A2B — all champions, 2026-08-31 / 2026-08-20 / 2026-08-05). Count is 3, but all in one persona.
    Personas present: champion. Missing: economic buyer, HR admin, IT security, finance.
    Most valuable add: economic buyer.
    On-file unengaged fit: none on file (no C-5A8FC2 rows).

NOT FLAGGED (3 deals, for completeness)
Deal-84DBA6 (63929535929): 3 active, 3 personas (champion, economic buyer, IT security).
Deal-4B0BEB (61038797752): 4 active, 4 personas (champion, economic buyer, HR admin, finance).
Deal-D348E1 (61750885954): 5 active, 5 personas (all five present).

ARITHMETIC SUMMARY
14 deals total. Flagged 11 (5 single-threaded + 6 under-threaded); passed 3. Active-contact math per deal shown above; the only exclusions were is_former=true (3 contacts: CT-F2C1AE, CT-40B45, CT-86B22F) and last-engaged before 2026-07-26 (2 contacts: CT-A902AE, CT-913581).

GAPS IN THE PROVIDED DATA
- No amount field → cannot report deal amounts.
- No stage field → cannot report stages, cannot verify "open," and persona recommendations are gap-based rather than stage-informed.
- No unengaged contacts on file for C-E23238, C-535D36, C-737030, C-5A8FC2 → "none on file" for those four flagged deals.
talk-track-from-transcriptsPASS 1.00
Plain-English question: From this rep's last 10 calls: how do they open, handle objections, and how often do they lock a next step?
Prompt — what the model saw
From this rep's last 10 calls: what they lead with in the first five minutes, how they handle the three most common objections, how often a concrete next step is agreed (rate), and every competitor a prospect raised. Quote at most one line per finding and only lines that appear in the transcripts. End with two coaching notes.
cost $0.0390 · 68s · in 1,694 / out 3,145 tok
✓ objection_count✓ next_step_rate✓ competitors✓ coaching✓ quote_used

✓ Correct Answer (ground truth)

Plain-English answer: 8 of 10 calls open with the same customer-win story; objection handling, next-step agreement rate, and every competitor raised, with quotes.
The answer in one line: n_transcripts = 10; objection_count = 3; rep = Alex Franklin (+3 more below)
n_transcripts10
objection_count3
allowed_new13 items — e.g. TT-001, TT-002, TT-003, TT-004
competitor_quotesAwardco=We're also in late talks with Awardco — their rewards catalog looks bigger than yours.; Kudos=How are you different from Kudos? Our CEO used them at her last company.
competitors_prospect_raisedAwardco, Kudos
next_stepagreed=7; total=10; rate=0.7
objectionsobjection=budget; count=4; quote=Honestly, budget is locked until next fiscal year — I can't add a new line item right now.; transcript_id=TT-001 | objection=timing; count=3; quote=This is good, but can we revisit it next quarter? Open enrollment starts in October and we're underwater.; transcript_id=TT-002 | objection=in-house; count=3; quote=We already do recognition with a spreadsheet and quarterly gift cards — why would we change?; transcript_id=TT-004
opener_themetheme=opens with a customer recognition win story (400-person retailer, 18% turnover reduction); count=8; total=10; quote=Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it.; transcript_id=TT-001
repAlex Franklin
rep_raised_competitors_excludedWorkhuman
graded checksobjection_count · next_step_rate · competitors · coaching · quote_used

Model's actual answer

show the model's full answer
ANALYSIS: ALEX FRANKLIN, LAST 10 CALLS (TT-001 through TT-010)

Note on data: transcripts contain only minutes 0–15, so all findings are limited to the lines provided.

1) WHAT THEY LEAD WITH (first five minutes)

8 of 10 calls (TT-001, -002, -003, -005, -006, -007, -008, -010) open with the identical proof-point story at minute 0:
"Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it."

2 of 10 open differently: TT-004 (Deal-403845) with an agenda ("I put together a short agenda — security review first, then pricing.") and TT-009 (Deal-1E2498) with pricing ("You asked for straight pricing last time, so let's start there."). Arithmetic: 8/10 = 80% story-led, 2/10 = 20% agenda/pricing-led.

2) THREE MOST COMMON OBJECTIONS AND HANDLING

a. Budget locked (4 of 10 calls: TT-001, -003, -006, -010). Handled identically every time by reframing cost as turnover savings:
"Totally fair. Most teams fund this out of turnover savings — that retailer saved about $210k in avoided backfills, which is how their finance team signed off."

b. Timing — revisit next quarter / open enrollment (3 of 10 calls: TT-002, -005, -008). Handled every time with a 90-day single-department pilot offer:
"Makes sense. What if we scope a 90-day pilot with one department so you have internal data before next quarter's planning?"

c. Status quo — spreadsheet + gift cards (3 of 10 calls: TT-004, -007, -009). Handled every time with automation + analytics differentiation:
"Spreadsheets work until they scale — the difference is automation: milestones fire without HR lifting a finger, and you get analytics on who is being recognized."

Arithmetic: 4 + 3 + 3 = 10 objection occurrences across 10 calls. Each response is verbatim-identical across its occurrences.

3) CONCRETE NEXT STEP AGREED — RATE

7 of 10 calls end with the prospect committing to a specific working session (TT-001, -002, -003, -005, -006, -008, -009):
"Yes, Thursday at 2pm works — send the invite and I'll bring our HRIS manager."

3 of 10 end with no next step: TT-004 (Deal-403845, committee), TT-007 (Deal-EDC141, "I need to think about it"), TT-010 (Deal-84DBA6, committee). Arithmetic: 7/10 = 70% next-step rate.

4) COMPETITORS RAISED BY PROSPECTS

Two competitors, each raised once:
- Awardco (TT-003, Deal-547B2B): "We're also in late talks with Awardco — their rewards catalog looks bigger than yours."
- Kudos (TT-007, Deal-EDC141): "How are you different from Kudos? Our CEO used them at her last company."

Excluded: Workhuman appears in TT-005 (Deal-C61CF7) but was raised by the rep, not the prospect, so it does not count per your question.

COACHING NOTES

1. The pitch is fully scripted — same opener in 8/10 calls and verbatim-identical responses to all three top objections. The story is strong, but when the prospect pushes past the script (committee stalls in TT-004 and TT-010, "no urgency" in TT-007), the rep has no second move and simply accepts the stall ("Understood — I'll leave it with you."; "Fair enough."). Build a planned response for stalls that don't match the top three.

2. In all 3 no-next-step calls the rep closed with zero commitment — no follow-up date, no smaller ask — while in all 7 successful calls the "next step" is the identical working-session ask at minute 14. Coach a fallback micro-commitment (e.g., a dated check-in) so stalled calls don't end at 0/3, and vary the ask so it fits the prospect rather than firing at the same minute every time.
reporting-analytics 4 tests
q3-weighted-forecastPASS 1.00
Plain-English question: Build the Q3 2026 weighted forecast from the open pipeline, and show the math deal by deal.
Prompt — what the model saw
Below are the open deals in the default sales pipeline with a close date between 2026-07-01 and 2026-10-15, pulled from the CRM warehouse on 2026-09-05. Stages: DS1 (discovery) through DS5 (contract). Build the Q3 2026 forecast (quarter = 2026-07-01 to 2026-09-30). Weighting rule: 100 percent of COMMIT plus 35 percent of BEST_CASE. PIPELINE deals count zero. Only deals with a close date inside the quarter count. Report: COMMIT total, BEST_CASE total, weighted forecast, count of deals in each category inside the quarter, the deals you excluded for being outside the quarter (count and total amount), and the top 5 BEST_CASE deals by amount inside the quarter. Show arithmetic. Then, in a final '## Data quality' section of no more than four sentences, name the data quality problems in this extract that would make the forecast unreliable if it were run unattended.
cost $0.1481 · 165s · in 27,510 / out 8,160 tok
✓ commit_total_44729✓ best_case_total_203565✓ weighted_115976.75✓ commit_count_7✓ best_case_count_24✓ excluded_count_32✓ excluded_amount_227575✓ top5_aliases✓ flags_blank_owner✓ flags_missing_why_buys

✓ Correct Answer (ground truth)

Plain-English answer: A stage-weighted Q3 forecast number computed from the open pipeline, with per-stage math that reconciles to the deal data.
The answer in one line: excluded_amount = 227575; best_case_total = 203565; weighted_forecast = 115,976.75; commit_total = 44729; excluded_out_of_quarter = 32; best_case_count = 24; flags = blank owner column + empty why_buys fields must be surfaced (+2 more below)
excluded_amount227,575
best_case_count24
best_case_total203,565
commit_count7
commit_total44,729
excluded_out_of_quarter32
weighted_forecast115,976.75
flagsblank owner column + empty why_buys fields must be surfaced
top5_best_caseDeal-2D7423, Deal-25F752, Deal-E53952, Deal-5EED42, Deal-FA32A0
graded checkscommit_total_44729 · best_case_total_203565 · weighted_115976.75 · commit_count_7 · best_case_count_24 · excluded_count_32 · excluded_amount_227575 · top5_aliases · flags_blank_owner · flags_missing_why_buys

Model's actual answer

show the model's full answer
Q3 2026 FORECAST (quarter = 2026-07-01 to 2026-09-30)

Extract: 86 open deals pulled 2026-09-05. 54 fall inside the quarter; 32 fall outside and are excluded.

COMMIT TOTAL (in quarter): 44,729
7 deals: Deal-547B2B (11,200) + Deal-B7EBD1 (9,000) + Deal-403845 (9,000) + Deal-A2B47C (6,360) + Deal-2465CE (5,400) + Deal-A5E80A (2,520) + Deal-499BF6 (1,249)
= 11,200 + 9,000 + 9,000 + 6,360 + 5,400 + 2,520 + 1,249 = 44,729

BEST_CASE TOTAL (in quarter): 203,565
24 deals: 38,935 + 24,000 + 19,656 + 16,250 + 11,116 + 10,800 + 10,500 + 9,890 + 9,720 + 9,000 + 7,200 + 3,840 + 3,780 + 3,600 + 3,240 + 3,150 + 3,120 + 3,060 + 2,916 + 2,760 + 2,484 + 2,100 + 1,920 + 528 = 203,565

WEIGHTED FORECAST:
100% x 44,729 + 35% x 203,565
= 44,729 + 71,247.75
= 115,976.75

DEAL COUNTS IN QUARTER:
- COMMIT: 7
- BEST_CASE: 24
- PIPELINE: 23 (total 201,637.40 — counts zero toward the forecast)

EXCLUDED — OUTSIDE QUARTER (close dates 2026-10-01 to 2026-10-15): 32 deals, 227,575 total
Breakdown: 1 COMMIT (Deal-D348E1, 13,770), 9 BEST_CASE (28,240: Deal-C61CF7 5,400 + Deal-48B656 5,160 + Deal-901332 3,600 + Deal-47AE31 3,600 + Deal-15D24F 3,600 + Deal-ED725A 2,400 + Deal-8AD4A5 1,800 + Deal-5FDCE4 1,600 + Deal-F5A622 1,080), 22 PIPELINE (185,565).
Check: 13,770 + 28,240 + 185,565 = 227,575.

TOP 5 BEST_CASE DEALS IN QUARTER (by amount):
1. Deal-2D7423 — 38,935 (close 2026-09-30, DS3)
2. Deal-25F752 — 24,000 (close 2026-09-25, DS4)
3. Deal-E53952 — 19,656 (close 2026-09-30, DS4)
4. Deal-5EED42 — 16,250 (close 2026-09-30, DS3)
5. Deal-FA32A0 — 11,116 (close 2026-09-25, DS3)
Top-5 sum: 38,935 + 24,000 + 19,656 + 16,250 + 11,116 = 109,957 — i.e., 54% of the BEST_CASE total rests on five deals.

## Data quality
Owner is blank on 85 of 86 deals (only Deal-C9C286 has an owner), so there is no rep-level accountability behind any forecast category. Two COMMIT deals sit in early stages (Deal-A5E80A at DS1, Deal-499BF6 at DS2), contradicting the COMMIT label, and four PIPELINE deals (Deal-31AD2C, Deal-333EBB, Deal-57FF13, Deal-7A2454) had close dates already past at the 2026-09-05 extract but are still open, indicating stale records. Close dates cluster heavily on 2026-09-30 (11 in-quarter deals), a classic quarter-end date-padding pattern that inflates the in-quarter population, and 71 of 86 deals have zero why_buys_chars, suggesting unqualified pipeline. One deal (Deal-42326B) carries a fractional amount (2,480.40) amid otherwise whole-dollar values, hinting at currency or data-entry inconsistencies.
aha-moment-2x20.83
Plain-English question: Do companies that hit 5+ givers and one redemption in month one actually retain better at 24 months?
Prompt — what the model saw
Test the activation hypothesis on the company cohort below: companies with 5+ unique givers (m1_users >= 5) AND 1+ successful redemption (m1_redemptions >= 1) in their first calendar month retain better at 24 months than those with one signal or neither. Every company in this extract is already 25+ months old, so current_status = 'active' means retained at 24 months. Report the full 2x2 (cohort size and 24-month retention rate for: both signals, givers-only, redemption-only, neither), state how many companies were excluded from the denominator and why (if none, say so), name the single signal with the largest retention lift, and state what this does and does not prove.
cost $0.1620 · 383s · in 8,760 / out 20,116 tok
✓ both_66✓ givers_47✓ redemp_31✓ neither_40✓ excluded_0✗ largest_lift_givers

✓ Correct Answer (ground truth)

Plain-English answer: The four cohort cells are graded against computed 24-month retention; the numbers come from the scorer's cohort table, not opinion.
The answer in one line: both_5plus_givers_and_redemption = 66.0% (31 of 47)
both_5plus_givers_and_redemption66.0% (31 of 47)
exclusionsnone — claim bounded as correlation, not causation
givers_only46.9% (23 of 49)
largest_single_signal_liftunique givers
neither40.0% (38 of 95)
redemption_only31.0% (9 of 29)
graded checksboth_66 · givers_47 · redemp_31 · neither_40 · excluded_0 · largest_lift_givers

Model's actual answer

show the model's full answer
# Activation Hypothesis Test — 2023 Signup Cohort Extract

**Definitions used (as given):** Givers signal = m1_users >= 5; redemption signal = m1_redemptions >= 1; retained at 24 months = current_status = 'active'. All other statuses (cancelled, non_renewing) = not retained.

## The 2x2

| Cell | n | Retained (active) | 24-mo retention | Arithmetic |
|---|---|---|---|---|
| Both signals | 47 | 31 | 66.0% | 31/47 = 0.6596 |
| Givers-only (users>=5, redemptions=0) | 49 | 23 | 46.9% | 23/49 = 0.4694 |
| Redemption-only (users<5, redemptions>=1) | 29 | 9 | 31.0% | 9/29 = 0.3103 |
| Neither | 95 | 38 | 40.0% | 38/95 = 0.4000 |
| **Total** | **220** | **101** | **45.9%** | 101/220 = 0.4591 |

Cell sizes sum to 220; retained counts sum to 101 — matches the file totals (101 active, 116 cancelled, 3 non_renewing).

## Exclusions

**None. 0 companies excluded; denominator = all 220.** Every row has non-missing m1_users, m1_redemptions, and current_status, and per the prompt all are 25+ months old. Two classification notes, not exclusions:
- 9 rows with m1_users = 0 were kept (zero is an observed value): 8 land in "neither"; C-F0A562FF (m1_users=0, m1_redemptions=10) lands in redemption-only — a data oddity (redemptions with no recorded m1 users/gives), classified per the rules as given.
- The 3 non_renewing companies (C-0B2078FB → neither, C-0A96134F and C-0BEAF685 → redemption-only) are counted as NOT retained per the "active = retained" rule. Sensitivity: if non_renewing were instead counted as retained, redemption-only would be 11/29 = 37.9% and neither 39/95 = 41.1%; the "both" cell and the ordering are unchanged.

## Largest single-signal lift

**The givers signal (m1_users >= 5): +6.9 percentage points** over neither (46.9% vs 40.0%). Redemption-only is *negative* lift: 31.0% vs 40.0% = **−9.0 pp** — a first-month redemption without 5+ users is associated with worse retention than no signal at all in this extract.

## Verdict on the hypothesis

**Supported directionally in this extract.** Both signals (66.0%) beat each single-signal cell and neither: +19.0 pp vs givers-only, +34.9 pp vs redemption-only, +26.0 pp vs neither. Full ordering: both > givers-only > neither > redemption-only.

## What this does and does not prove

**Does:** In these 220 companies, the combined month-1 pattern (5+ users AND 1+ redemption) is strongly associated with higher 24-month active status, and the two signals are not additive equals — the user-count signal carries the single-signal lift while redemption-alone correlates with churn.

**Does not:**
- **Causation.** This is observational with no controls. Confounding is visible in the extract itself: retained companies are overwhelmingly tier_three with nonzero current_arr_usd, while churned ones are overwhelmingly tier_two with $0 ARR. Tier/monetization path is entangled with both activation capacity and retention; activation may be a marker of healthy accounts, not a driver.
- **The stated signal exactly.** "5+ unique givers" was proxied by m1_users >= 5 per your operationalization. m1_users is unique users, not unique givers (e.g., C-0DD506E9: 27 m1_users but only 2 m1_gives). The extract contains no unique-givers field, so the hypothesis as literally worded ("givers") is not directly testable with these columns.
- **Statistical reliability.** No significance testing; cells are small (both n=47, redemption-only n=29), and this is a single extract of unknown sampling — rates carry wide implicit error bars.
- **Mechanism.** Nothing here shows *when* companies churned (no intermediate status history), whether "first calendar month" is tenure-normalized for mid-month signups, or that driving the activation pattern would change outcomes rather than select for already-durable accounts.
arr-rollforward-reconciliationPASS 1.00
Plain-English question: Does CRM ARR match billing ARR — and if not, exactly where is the variance?
Prompt — what the model saw
Reconcile ARR between CRM company records and billing subscriptions as of 2026-09-05. Billing ARR is MRR x 12. Report both totals, the variance, and a decomposition into buckets (status mismatch, rounding, missing records, other) that sums exactly to the variance. List each mismatched account with a suggested owner. Business rule: any subscription with a term other than 12 months must have cf_agreement_end_date populated; list violations.
cost $0.1035 · 309s · in 561 / out 14,459 tok
✓ billing_total✓ crm_total✓ variance✓ buckets_sum✓ violations_listed

✓ Correct Answer (ground truth)

Plain-English answer: Billing $604,739.28 vs CRM $603,581.76: a $1,157.52 variance decomposed into status-mismatch, rounding, and missing-record buckets.
The answer in one line: n_companies = 40; n_mismatched_accounts = 7; as_of = 2026-09-05 (+4 more below)
n_companies40
n_mismatched_accounts7
as_of2026-09-05
billing_arr_total604739.28
bucket_sum_equals_variance1
buckets{"status_mismatch": {"amount": "-13158.48", "accounts": [{"company_alias": "C-0C8323BF", "billing_arr": "0.00", "hubspot_arr": "4905.24", "diff": "-4905.24", "suggested_owner": "Dana Mercer"}, {"compa
crm_arr_total603581.76
ruleterm_months != 12 requires cf_agreement_end_date
variance1157.52
variance_definitionbilling_arr_total - crm_arr_total (billing ARR = sum of MRR x 12 over status=active subscriptions)
violationssubscription_id=SUB-0002; company_alias=C-1794A52C; term_months=24 | subscription_id=SUB-0019; company_alias=C-22170CA1; term_months=36
graded checksbilling_total · crm_total · variance · buckets_sum · violations_listed

Model's actual answer

show the model's full answer
ARR RECONCILIATION — AS OF 2026-09-05
Scope: 39 billing subscriptions, 39 CRM company records. Billing ARR = MRR x 12. Cancelled subscriptions are excluded from the billing ARR total (a cancelled sub generates no billable ARR); the ARR those accounts still carry in CRM is captured in the status-mismatch bucket.

TOTALS
- Billing ARR (37 active subs): $604,739.28
  Arithmetic: sum of active MRR = $50,394.94; x 12 = $604,739.28
  (The 2 cancelled subs, SUB-000E + SUB-000F, MRR $408.77 + $687.77 = $1,096.54, ARR $13,158.48, are excluded. If included, all-sub billing ARR = $617,897.76.)
- CRM ARR (39 company records): $603,581.76
  Arithmetic: sum of hubspot_arr across all 39 records = $603,581.76
- Variance (CRM minus Billing): $603,581.76 - $604,739.28 = -$1,157.52
  CRM understates billing ARR by $1,157.52.

DECOMPOSITION (sums exactly to -$1,157.52)
1. Status mismatch: +$13,158.48
   CRM carries full ARR for 2 cancelled subscriptions:
   - C-0C8323BF: $408.77 x 12 = $4,905.24 (sub SUB-000E cancelled)
   - C-0DC4FB8C: $687.77 x 12 = $8,253.24 (sub SUB-000F cancelled)
   $4,905.24 + $8,253.24 = $13,158.48

2. Missing records: -$11,952.00 (net)
   - C-21629AA4: active sub SUB-0004, $2,370.77 x 12 = $28,449.24, but NO CRM record: -$28,449.24
   - C-0D5BBE3A: CRM record shows $16,497.24, but NO subscription in billing: +$16,497.24
   Net: -$28,449.24 + $16,497.24 = -$11,952.00

3. Rounding: +$36.00
   - C-0D66DF9E: billing $1,932.00 x 12 = $23,184.00; CRM $23,200.00; diff +$16.00
   - C-14D70CE0: billing $1,515.00 x 12 = $18,180.00; CRM $18,200.00; diff +$20.00
   Both CRM values are multiples of $100 while exact ARR is not ($23,184 -> $23,200; $18,180 -> $18,200), consistent with nearest-$100 rounding in CRM. This is an inference from the pattern; the data does not state the cause.

4. Other: -$2,400.00
   - C-0F7269D7: billing $2,233.00 x 12 = $26,796.00; CRM $24,396.00; diff -$2,400.00
   Note: $24,396.00 = $2,033.00 x 12, i.e., the CRM value implies an MRR $200/mo lower than billing. Root cause is not determinable from the data provided.

Bucket check: +$13,158.48 - $11,952.00 + $36.00 - $2,400.00 = -$1,157.52 (matches variance exactly)

All other accounts reconcile exactly (33 active accounts match to the cent).

MISMATCHED ACCOUNTS AND SUGGESTED OWNERS
The provided data contains no owner field, so no named owners can be cited. Suggestions below are functional roles only:
- C-0C8323BF (cancelled sub, CRM ARR $4,905.24) -> Billing Ops / subscription-status sync owner
- C-0DC4FB8C (cancelled sub, CRM ARR $8,253.24) -> Billing Ops / subscription-status sync owner
- C-21629AA4 (active sub $28,449.24, no CRM record) -> CRM administrator (create/restore company record)
- C-0D5BBE3A (CRM ARR $16,497.24, no subscription) -> Billing Ops (verify whether a subscription should exist)
- C-0D66DF9E (+$16.00 rounding) -> RevOps data steward
- C-14D70CE0 (+$20.00 rounding) -> RevOps data steward
- C-0F7269D7 (-$2,400.00 unexplained MRR delta) -> Billing Ops, jointly with the account owner (name not available in data)

BUSINESS RULE VIOLATIONS (term != 12 months with cf_agreement_end_date empty)
- SUB-0002 / C-1794A52C: term 24 months, cf_agreement_end_date blank -> VIOLATION
- SUB-0019 / C-22170CA1: term 36 months, cf_agreement_end_date blank -> VIOLATION
Compliant multi-term subs for reference: SUB-000C / C-0DB48281 (24 mo, end date 2027-11-30) and SUB-001A / C-0FC4DBB8 (36 mo, end date 2027-11-30).
Suggested owner for violations: Billing admin (populate agreement end dates).

DATA GAPS (explicit)
- No owner names/fields in either file; owner suggestions are role-based only.
- Cause of the C-0F7269D7 $2,400 delta and the C-0D5BBE3A orphan CRM record cannot be determined from the data provided.
- Currency is not stated; figures are reported as given.
kvm-month-over-monthPASS 1.00
Plain-English question: How did the four core KVMs move in August versus July?
Prompt — what the model saw
Report the four core KVMs for 2026-08 against 2026-07: giving rate, redemptions per user, 1:1 meetings engagement, pulse check engagement. For each: value, prior value, absolute and relative change, direction. Name the KVM with the largest relative move and the segment (plan_tier or size_band) driving it if the data supports that; otherwise say the data does not. Under 200 words plus one table.
cost $0.1915 · 480s · in 6,265 / out 23,614 tok
✓ four_kvms✓ biggest_mover_named✓ direction_correct✓ segment_named✓ word_cap

✓ Correct Answer (ground truth)

Plain-English answer: Giving rate 60.27% (up 0.07%), redemptions per user 1.7302 (up 0.01%), with 1:1 meetings and pulse engagement each up fractionally — direction reported for all four.
The answer in one line: n_company_rows_per_month = 30; aggregation = simple mean of company-level values per month
n_company_rows_per_month30
aggregationsimple mean of company-level values per month
biggest_moverpulse_engagement
driving_segmentsize_band=enterprise; 2026-07=0.55; 2026-08=0.2743; rel_change_pct=-50.13
kvms{"giving_rate": {"2026-07": 0.6023, "2026-08": 0.6027, "abs_change": 0.0004, "rel_change_pct": 0.07, "direction": "up"}, "redemptions_per_user": {"2026-07": 1.73, "2026-08": 1.7302, "abs_change": 0.00
segment_detail{"smb": {"2026-07": 0.6588, "2026-08": 0.6573, "rel_change_pct": -0.23}, "mid_market": {"2026-07": 0.593, "2026-08": 0.5942, "rel_change_pct": 0.2}, "enterprise": {"2026-07": 0.55, "2026-08": 0.2743,
graded checksfour_kvms · biggest_mover_named · direction_correct · segment_named · word_cap

Model's actual answer

show the model's full answer
**Core KVMs — 2026-08 vs 2026-07 (unweighted mean across 30 companies)**

| KVM | Aug 2026 | Jul 2026 | Abs change | Rel change | Direction |
|---|---|---|---|---|---|
| Giving rate | 0.6027 | 0.6023 | +0.0004 | +0.07% | Up |
| Redemptions per user | 1.7302 | 1.7300 | +0.0002 | +0.01% | Up |
| 1:1 meetings engagement | 0.4472 | 0.4469 | +0.0003 | +0.06% | Up |
| Pulse check engagement | 0.5086 | 0.6006 | −0.0920 | −15.31% | Down |

Giving rate, redemptions per user, and 1:1 meetings engagement were essentially flat (all within ±0.1% month-over-month).

Pulse check engagement shows the largest relative move: −15.31%. The enterprise size_band drives this — enterprise pulse engagement fell from 0.5500 to 0.2743 (−50.13%), with all ten enterprise companies dropping by roughly half (from ~0.53–0.57 to ~0.26–0.29). SMB (−0.22%) and mid-market (+0.21%) pulse engagement were flat. Plan_tier does not differentiate: all 30 companies are tier_three.
customer-success 5 tests
churn-save-eligibilityPASS 1.00
Plain-English question: Which at-risk accounts actually qualify for a churn-save offer, and which play fits each?
Prompt — what the model saw
Which at-risk accounts qualify for a churn-save offer under the documented eligibility rules, what amount is at stake per account and in total, and which play fits each (usage revival, executive touch, commercial concession)? Cite the signal that justifies each play. List accounts that look at risk but do not qualify and why.
cost $0.0936 · 265s · in 278 / out 12,957 tok
✓ eligible_set✓ total_at_stake✓ plays_cited✓ noneligible_named✓ rules_applied

✓ Correct Answer (ground truth)

Plain-English answer: Only accounts meeting all three rules qualify — e.g. C-0F6C0F34 at $49,707 (executive touch) and C-0B827671 at $25,365 (usage revival) — with amount and play per account.
The answer in one line: n_accounts = 30; snapshot_date = 2026-09-05
n_accounts30
eligibleaccount_alias=C-0F6C0F34; amount_at_stake=49707.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0B827671; amount_at_stake=25365.00; play=usage revival; justifying_signal=usage_trend_3m=declining over the last 3 months | account_alias=C-0B360C78; amount_at_stake=35748.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 75% - risk is commercial, not adoption | account_alias=C-0B0F1BAB; amount_at_stake=5494.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0CA21961; amount_at_stake=16829.00; play=usage revival; justifying_signal=seat utilization 26% is below 50% | account_alias=C-0E9C27D1; amount_at_stake=41235.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 85% - risk is commercial, not adoption … (+2 more)
non_eligible_at_riskaccount_alias=C-0BC71BDD; health_score=55 | account_alias=C-0BA71F12; health_score=52 | account_alias=C-0F6694C3; health_score=43 | account_alias=C-0BE96399; health_score=54 | account_alias=C-0F876796; health_score=47 | account_alias=C-0FCCD2DF; health_score=43 … (+1 more)
ruleshealth_score < 60, churn_save_eligible_amount > 0, renewal within 120 days of snapshot
snapshot_date2026-09-05
total_amount_at_stake224601.00
graded checkseligible_set · total_at_stake · plays_cited · noneligible_named · rules_applied

Model's actual answer

show the model's full answer
ELIGIBILITY SCREEN

Rules applied (all three must hold): R1 health_score < 60; R2 churn_save_eligible_amount > 0; R3 renewal_date within 120 days of snapshot 2026-09-05, i.e. on or before 2027-01-03 (25 remaining Sept days + 31 Oct + 30 Nov + 31 Dec + 3 Jan = 120).

15 of 30 accounts have health_score < 60. Of those, 8 pass all three rules.

QUALIFYING ACCOUNTS — TOTAL AT STAKE: $224,601.00

1. C-0F6C0F34 — $49,707.00 (ARR $86,741; renewal 2026-10-03, 28 days out)
   Play: Executive touch. Signal: champion_active=false — no active sponsor on the largest at-risk ARR; usage_trend_3m=growing at 78% seat utilization (308/395), so adoption is not the gap.

2. C-0B827671 — $25,365.00 (ARR $72,088; renewal 2026-11-14, 70 days)
   Play: Usage revival. Signal: usage_trend_3m=declining; seat utilization 113/202 = 56%.

3. C-0B360C78 — $35,748.00 (ARR $60,427; renewal 2026-10-28, 53 days)
   Play: Commercial concession. Signal: health_score=57 despite growing usage and an active champion at 75% utilization (246/327) — adoption and relationship signals are positive, so the residual risk is commercial; eligible amount = 59% of ARR (35,748/60,427).

4. C-0B0F1BAB — $5,494.00 (ARR $15,391; renewal 2026-09-23, 18 days)
   Play: Executive touch. Signal: champion_active=false plus the lowest qualifier health_score=38; most imminent renewal in the set.

5. C-0CA21961 — $16,829.00 (ARR $31,501; renewal 2026-12-28, 114 days)
   Play: Commercial concession. Signal: seats_used=84 of 325 = 26% utilization with flat usage — ~241 paid seats unused; right-sizing/price relief closes the value gap.

6. C-0E9C27D1 — $41,235.00 (ARR $75,093; renewal 2026-09-24, 19 days)
   Play: Commercial concession. Signal: renewal 19 days out with health_score=39 while utilization is already 85% (134/157) and champion_active=true — usage and relationship levers look healthy, leaving commercial terms as the fast lever.

7. C-0CEF69FD — $32,621.00 (ARR $79,324; renewal 2026-11-21, 77 days)
   Play: Executive touch. Signal: champion_active=false; usage_trend_3m=growing at 71% utilization (97/136).

8. C-0D3278C7 — $17,602.00 (ARR $33,815; renewal 2026-11-12, 68 days)
   Play: Usage revival. Signal: usage_trend_3m=declining; seat utilization 126/380 = 33%.

Arithmetic: 49,707 + 25,365 + 35,748 + 5,494 + 16,829 + 41,235 + 32,621 + 17,602 = $224,601.00 (= 49.4% of the qualifiers' combined ARR of $454,380).

Play mix: executive touch $87,822 (49,707+5,494+32,621); commercial concession $93,812 (35,748+16,829+41,235); usage revival $42,967 (25,365+17,602). Sums to 224,601.

AT RISK (HEALTH < 60) BUT NOT QUALIFYING

Fail R2 only (churn_save_eligible_amount = 0.00, renewal inside window):
- C-0BC71BDD — health 55, renewal 2026-10-27 (52 days). No save budget allocated. Risk signals: champion_active=false, flat usage, 59/197 = 30% utilization.
- C-0BE96399 — health 54, renewal 2026-10-29 (54 days). Eligible amount 0.00. Risk signals: declining usage, 43/154 = 28% utilization.
- C-10A56B0F — health 54, renewal 2026-12-12 (98 days). Eligible amount 0.00. Risk signals: declining usage, champion_active=false, 85/176 = 48% utilization.

Fail R3 only (eligible amount > 0, renewal beyond 120 days):
- C-0BA71F12 — health 52, eligible $6,824.00; renewal 2027-04-11 = 218 days out.
- C-0F876796 — health 47, eligible $19,958.00; renewal 2027-02-06 = 154 days out.
These two carry $26,782.00 of eligible amount that is outside the renewal window.

Fail both R2 and R3:
- C-0F6694C3 — health 43, eligible 0.00; renewal 2027-03-21 = 197 days out.
- C-0FCCD2DF — health 43, eligible 0.00; renewal 2027-04-23 = 230 days out.

Also: the 15 accounts with health_score >= 60 fail R1 outright (and all have churn_save_eligible_amount = 0.00 as well), so none are candidates.

DATA GAP (explicit): the files document eligibility rules only. No definitions of "usage revival," "executive touch," or "commercial concession," and no signal-to-play mapping, were provided. Play assignments above are analytical judgments derived strictly from the fields present (champion_active, usage_trend_3m, seats vs seats_used, health_score, renewal proximity) — they are not rule-based outputs.
expansion-kitPASS 1.00
Plain-English question: Build the expansion kit for the customer who just replied to our outreach.
Prompt — what the model saw
A customer replied to outreach (reply included). Build the expansion kit: seat coverage (licensed over headcount), two lines on usage health, headroom in seats and ARR at the current per-seat rate, who replied and whether they can buy, the right buyer from our contacts, and a reply email under 150 words that moves toward the right person without being pushy and cites one real usage fact.
cost $0.0221 · 31s · in 0 / out 1,241 tok
✓ coverage_pct✓ headroom_arr✓ right_buyer✓ replier_not_buyer✓ usage_fact✓ email_under_150

✓ Correct Answer (ground truth)

Plain-English answer: 37.5% seat coverage, 250 headroom seats worth $15,000 ARR at $60 a seat, usage up 43.2% in six months — and the replier cannot buy; Dana R. (VP People) is the right buyer.
The answer in one line: headroom_seats = 250; coverage_pct = 37.5; company_alias = C-0DDFC9A7 (+3 more below)
headroom_seats250
coverage_pct37.5
company_aliasC-0DDFC9A7
headroom_arr15000.00
per_seat_annual_rate60.00
replierMaria S.
replier_can_buy0
right_buyername=Dana R.; title=VP People
usage_factsactive_users_2026_08=126; active_users_2026_03=88; growth_pct_6m=43.2; licensed_seat_utilization_pct=84
usage_health_linesMonthly active users grew from 88 (2026-03) to 126 (2026-08), +43.2% over 6 months., Last month 126 of 150 licensed seats were active (84.0% utilization), approaching the seat cap.
graded checkscoverage_pct · headroom_arr · right_buyer · replier_not_buyer · usage_fact · email_under_150

Model's actual answer

show the model's full answer
EXPANSION KIT — C-0DDFC9A7

1) SEAT COVERAGE
Licensed seats: 150. Headcount: 400.
Coverage = 150 / 400 = 37.5% of employees licensed.
Unlicensed employees: 400 - 150 = 250.

2) USAGE HEALTH (two lines)
Monthly active users grew every month, 88 (Mar 2026) → 126 (Aug 2026): +38 users, ~43% growth over 5 months, ~8/month steady climb.
Seat utilization is 126 / 150 = 84% of licensed seats — adoption is real and approaching license capacity, not idle shelfware.

3) HEADROOM
Per-seat rate = current ARR / licensed seats = $9,000 / 150 = $60/seat/year.
Seat headroom = 400 - 150 = 250 unlicensed employees.
ARR headroom = 250 × $60 = $15,000 potential expansion ARR at current rate.
(For reference: unused licensed seats today = 150 - 126 = 24.)

4) WHO REPLIED / CAN SHE BUY
Maria S., People Operations Coordinator. No — she states purchasing and budget sit with Dana R. She has offered an introduction.

5) RIGHT BUYER
Dana R., VP People — owns budget and seat expansion per Maria's reply, and has been asking about usage numbers. Last engaged 2026-05-18, so a warm intro from Maria (engaged 2026-09-02) is the best path. Sam K. (Office Manager, last engaged 2025-11-03) is not the buyer; engagement is stale.

6) REPLY EMAIL (~100 words)

Subject: RE: Growing your team's recognition program

Hi Maria,

Thanks for the kind words — great to hear the recognition feed is busy every week.

Happy to help with the usage numbers Dana's been asking about. Since March, your monthly active users have grown from 88 to 126 — up about 43% — and you're now using 84 of every 100 licensed seats. I can put together a short summary she'd find useful for planning.

If an introduction makes sense, I'd welcome it — otherwise feel free to forward this along. Whichever is easiest for you.

Thanks again,
Cole

DATA GAPS (explicit): No data on Dana R.'s current sentiment beyond Maria's secondhand comment; no contact seniority/budget-authority data beyond titles; no renewal date or contract terms provided.
onboarding-statusPASS 1.00
Plain-English question: What is done and what is missing in this customer's onboarding before the mid-onboarding call?
Prompt — what the model saw
Prep the CSM for the mid-onboarding call: what is complete (integrations connected, allowance set, admins added, first recognition given, first redemption), what is not, early engagement signals, and the three things to cover on the call. Mark nothing complete without a data field that shows it.
cost $0.0285 · 48s · in 0 / out 2,317 tok
✓ complete_items✓ incomplete_items✓ three_items✓ grounded

✓ Correct Answer (ground truth)

Plain-English answer: 25 days live: Slack connected, allowance set, admins added — but the HRIS integration is not connected, with early engagement signals and three call priorities.
The answer in one line: days_live = 25; company_alias = C-0D284E42
days_live25
call_itemsConnect the HRIS integration (integration_hris is empty) to automate user provisioning and offboarding., Get to first redemption (first_redemption_at is empty): walk admins through the rewards catalog and confirm allowance funding., Build the giving habit: daily active givers averaged 4.3 in week 1 vs 13.0 last week - agree on manager nudges and a recognition cadence.
checklist{"slack_integration_connected": {"complete": true, "field": "integration_slack", "value": "2026-08-12"}, "hris_integration_connected": {"complete": false, "field": "integration_hris", "value": ""}, "a
company_aliasC-0D284E42
engagement_signalsavg_active_givers_week1=4.3; avg_active_givers_last_week=13; latest_active_givers=15; days_with_activity=25
graded checkscomplete_items · incomplete_items · three_items · grounded

Model's actual answer

show the model's full answer
CSM CALL PREP — C-0D284E42 (mid-onboarding, ~24 days post-signup)

WHAT IS COMPLETE (each backed by a data field)
- Slack integration connected — integration_slack = 2026-08-12 (1 day after signup 2026-08-11)
- Allowance set — allowance_set = 2026-08-13 (2 days after signup)
- Admins added — admins_added = 2
- First recognition given — first_recognition_at = 2026-08-15 14:22 (4 days after signup)

WHAT IS NOT COMPLETE
- HRIS integration — integration_hris field is empty; no date recorded
- First redemption — first_redemption_at field is empty; no date recorded

EARLY ENGAGEMENT SIGNALS (from onboarding_usage.csv, 2026-08-11 → 2026-09-04)
- Active givers grew 3 → 15: 15 − 3 = 12 net new; 15 / 3 = 5.0x over 25 days
- No zero days: minimum daily active givers = 3
- Weekly averages: Week 1 (Aug 11–17): (3+3+4+4+5+4+7) = 30 / 7 ≈ 4.3 | Week 2 (Aug 18–24): (5+7+6+9+8+9+9) = 53 / 7 ≈ 7.6 | Week 3 (Aug 25–31): (9+11+10+10+11+13+11) = 75 / 7 ≈ 10.7 | Week 4 partial (Sep 1–4): (13+13+15+15) = 56 / 4 = 14.0
- Last 7 days (Aug 29–Sep 4) average: (11+13+11+13+13+15+15) = 91 / 7 = 13.0 — roughly 3x the Week 1 average (13.0 / 4.3 ≈ 3.0x)
- Peak = 15 active givers on both 2026-09-03 and 2026-09-04; the trend is still rising, not plateaued

DATA GAPS (stated explicitly)
- Total employee headcount is not provided, so adoption rate (% of company actively giving) cannot be computed
- No data on reward catalog status, redemption attempts, or why redemption hasn't occurred — only that first_redemption_at is empty
- No data beyond admin count (2); admin identities/roles not provided

THREE THINGS TO COVER ON THE CALL
1. HRIS integration — the empty integration_hris field is the biggest structural gap. Ask what's blocking it and set a target date; until it's connected, user sync is manual.
2. First redemption — recognition has been flowing for ~3 weeks (first_recognition_at = 2026-08-15) but no one has redeemed (first_redemption_at empty). Ask whether users can see/access the reward catalog and know how to redeem. Do not assume a cause — the data only shows the redemption hasn't happened, not why.
3. Sustaining the growth curve — active givers 5x'd (3 → 15) with rising weekly averages (4.3 → 7.6 → 10.7 → 14.0). Ask what's driving it and how to keep it going; since headcount is unknown, whether 15 is 10% or 90% of the company cannot be determined from this data.
renewal-risk-conflicting-datesPASS 1.00
Plain-English question: Two systems disagree on renewal dates — which do you trust per account, and who is actually renewing inside 90 days?
Prompt — what the model saw
Produce a 90-day renewal risk brief. Two systems hold renewal dates and disagree on some accounts; decide which to trust per account and say why (multi-year contracts are known to be wrong in ChurnZero). For every renewal: company, CSM, ARR, date used, seat utilization, 3-month usage trend, risk rating with one sentence of evidence. Flag every disagreement. Close with total ARR renewing and ARR at risk.
cost $0.2153 · 466s · in 13,596 / out 23,826 tok
✓ total_renewing✓ arr_at_risk✓ disagreements_flagged✓ trust_rule

✓ Correct Answer (ground truth)

Plain-English answer: Chargebee wins per account (multi-year contracts are known wrong in ChurnZero), with each trusted date, the reason, and a 90-day-window flag.
The answer in one line: n_accounts = 20; n_disagreements = 5; snapshot_date = 2026-09-05 (+2 more below)
n_accounts20
n_disagreements5
accounts20 items — e.g. account_alias=C-0B144C78; csm=Cole Ingram; arr=30899.00; trusted_renewal_date=2026-11-02; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.4; usage_3m_ratio=1.03; risk=low; evidence=3-month usage ratio 1.03 (last3 avg 103 vs prior3 100), seat utilization 75% | account_alias=C-0B20DB64; csm=Dana Mercer; arr=21770.00; trusted_renewal_date=2026-10-07; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=56.6; usage_3m_ratio=1; risk=medium; evidence=3-month usage ratio 1.00 (last3 avg 295 vs prior3 295), seat utilization 57% | account_alias=C-0B344485; csm=Elena Sinclair; arr=64384.00; trusted_renewal_date=2026-11-16; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=78; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 241 vs prior3 231), seat utilization 78% | account_alias=C-0B369871; csm=Dana Mercer; arr=85128.00; trusted_renewal_date=2026-10-29; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.1; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 330 vs prior3 316), seat utilization 75%
arr_at_risk_90d359409.00
disagreementsaccount_alias=C-0B7D2C30; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-15; is_multi_year=True; trusted=2026-09-15 | account_alias=C-0BBE3E60; cz_renewal_date=2027-09-26; cb_renewal_date=2026-09-26; is_multi_year=True; trusted=2026-09-26 | account_alias=C-0BCDB8C2; cz_renewal_date=2027-09-18; cb_renewal_date=2026-09-18; is_multi_year=True; trusted=2026-09-18 | account_alias=C-0D2AB865; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-22; is_multi_year=True; trusted=2026-09-22 | account_alias=C-0F5D2323; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-29; is_multi_year=True; trusted=2026-09-29
snapshot_date2026-09-05
total_arr_renewing_90d1048715.00
trust_rulemulti-year contracts: Chargebee is authoritative (ChurnZero known wrong); otherwise systems agree or Chargebee wins
window2026-09-05 to 2026-12-04
graded checkstotal_renewing · arr_at_risk · disagreements_flagged · trust_rule

Model's actual answer

show the model's full answer
90-DAY RENEWAL RISK BRIEF — window 2026-09-24 to 2026-12-23
All 20 accounts have a renewal date in or before the window (3 are already past due as of 2026-09-24). No data is missing: all 20 accounts have complete CZ, Chargebee, and 12-month usage records.

METHOD
- Date source: For the 5 multi-year contracts, ChurnZero is known-wrong, so I use the Chargebee date. For the 15 single-year contracts, both systems agree exactly, so no decision was needed.
- Seat utilization = seats_used / seats. 3-month trend = active users Jun→Aug 2026, % change = (Aug − Jun) / Jun.
- Risk rubric (stated explicitly since none was given): High = 3-mo usage decline >5% OR utilization <35%; Medium = utilization 35–70% with no High trigger; Low = utilization ≥70% and stable/growing usage.

RENEWALS (sorted by date used; util = seats_used/seats; trend = Jun→Aug active users)

Company      CSM            ARR        Date used   Src  Util        3-mo trend    Risk
C-0B7D2C30   Dana Mercer    65,901     2026-09-15  CB   57.6%       97→84 (−13%)  High  *past due*
C-0BCDB8C2   Cole Ingram    54,427     2026-09-18  CB   54.7%      127→110 (−13%)  High  *past due*
C-0D2AB865   Elena Sinclair 38,022     2026-09-22  CB   61.4%      125→109 (−13%)  High  *past due*
C-0BBE3E60   Dana Mercer    30,993     2026-09-26  CB   64.9%       39→33 (−15%)  High
C-0F5D2323   Cole Ingram    90,647     2026-09-29  CB   28.5%       20→18 (−10%)  High
C-0EC6999D   Elena Sinclair 79,419     2026-10-03  both 27.7%       17→15 (−12%)  High
C-0B20DB64   Dana Mercer    21,770     2026-10-07  both 56.6%     294→294 ( 0%)  Medium
C-0BBC4E7A   Cole Ingram    56,374     2026-10-10  both 67.7%     142→139 (−2%)  Medium
C-0FD551AB   Elena Sinclair 48,815     2026-10-14  both 55.9%     123→126 (+2%)  Medium
C-0F9F8F13   Dana Mercer    46,230     2026-10-18  both 56.5%     185→182 (−2%)  Medium
C-0BC34584   Cole Ingram    16,740     2026-10-22  both 66.2%     104→106 (+2%)  Medium
C-0B7A7546   Elena Sinclair 35,062     2026-10-25  both 88.8%      64→63 (−2%)   Low
C-0B369871   Dana Mercer    85,128     2026-10-29  both 75.1%    326→333 (+2%)   Low
C-0B144C78   Cole Ingram    30,899     2026-11-02  both 75.4%    101→106 (+5%)   Low
C-0FC4DBB8   Elena Sinclair 94,732     2026-11-05  both 76.7%    189→193 (+2%)   Low
C-0D5BBE3A   Dana Mercer    39,740     2026-11-09  both 83.3%     88→91 (+3%)    Low
C-0FB9D5AF   Cole Ingram    63,158     2026-11-13  both 72.4%    173→176 (+2%)   Low
C-0B344485   Elena Sinclair 64,384     2026-11-16  both 78.0%    238→244 (+3%)   Low
C-0CB2C1B4   Dana Mercer    40,628     2026-11-20  both 81.6%     47→49 (+4%)    Low
C-22170CA1   Cole Ingram    45,646     2026-11-24  both 85.4%    143→146 (+2%)   Low

DATE DISAGREEMENTS (all 5 are multi-year contracts; Chargebee trusted per the known ChurnZero multi-year defect)
1. C-0B7D2C30: CZ 2026-09-10 vs CB 2026-09-15 → used CB 09-15 (now past due)
2. C-0BCDB8C2: CZ 2027-09-18 vs CB 2026-09-18 → used CB; CZ appears to show end-of-term (2027) instead of the annual renewal (2026)
3. C-0D2AB865: CZ 2026-09-10 vs CB 2026-09-22 → used CB 09-22 (now past due)
4. C-0BBE3E60: CZ 2027-09-26 vs CB 2026-09-26 → used CB; same end-of-term pattern as #2
5. C-0F5D2323: CZ 2026-09-10 vs CB 2026-09-29 → used CB 09-29
Note: CZ shows the identical date 2026-09-10 for three different multi-year accounts, a further data-quality signal supporting Chargebee for these. The 15 single-year accounts match exactly between systems — no disagreements.

RISK EVIDENCE (one sentence per account)

HIGH
- C-0B7D2C30: Active users fell 155→84 over 12 months (−46%) with −13% in the last 3, and the 09-15 renewal is already past due.
- C-0BCDB8C2: Usage dropped 200→110 over 12 months (−45%) and the 09-18 renewal is already past due.
- C-0D2AB865: Usage fell 199→109 over 12 months (−45%) and the 09-22 renewal is already past due.
- C-0BBE3E60: Steepest 3-month decline in the book (39→33, −15%) on top of a 12-month drop of 63→33 (−48%).
- C-0F5D2323: Only 28.5% seat utilization (111/390) with usage down 10% in 3 months — largest ARR in the high tier at $90,647.
- C-0EC6999D: 27.7% seat utilization (31/112) and usage drifting down (17→15, −12%) despite flat 12-month totals.

MEDIUM
- C-0B20DB64: Usage is stable (293→294 over 12 months) but 43% of seats (164/378) are unused.
- C-0BBC4E7A: Flat usage (142→139, −2% over 12 months) with a third of seats (109/337) idle.
- C-0FD551AB: Slightly growing usage (+2% in 3 months) but 44% of seats (166/376) unused.
- C-0F9F8F13: Flat usage (182→182 over 12 months) with 43.5% of seats (153/352) idle — downsizing exposure.
- C-0BC34584: Stable-to-growing usage (+3% over 12 months) but 33.8% of seats (167/494) unused.

LOW
- C-0B7A7546: 88.8% utilization (182/205) and usage up 8.6% over 12 months (58→63).
- C-0B369871: Usage grew every month, 289→333 (+15%), at 75.1% utilization (317/422).
- C-0B144C78: Usage up 17.8% over 12 months (90→106) at 75.4% utilization (169/224).
- C-0FC4DBB8: Largest account in the book ($94,732) growing steadily, 168→193 (+15%), 76.7% utilization (356/464).
- C-0D5BBE3A: Usage up 19.7% over 12 months (76→91) at 83.3% utilization (85/102).
- C-0FB9D5AF: Usage up 14.3% over 12 months (154→176) at 72.4% utilization (144/199).
- C-0B344485: Usage up 15.6% over 12 months (211→244) at 78.0% utilization (224/287).
- C-0CB2C1B4: Usage up 14.0% over 12 months (43→49) at 81.6% utilization (386/473).
- C-22170CA1: Usage up 12.3% over 12 months (130→146) at 85.4% utilization (251/294).

TOTALS
- Total ARR renewing in window: $1,048,715.00 (all 20 accounts; $158,350.00 of it — C-0B7D2C30, C-0BCDB8C2, C-0D2AB865 — is already past its renewal date as of 2026-09-24).
- ARR at risk (High): $359,409.00 (34.3% of renewing ARR).
- ARR on watch (Medium): $189,929.00 (18.1%).
- Combined High + Medium: $549,338.00 (52.4%).
- Low risk: $499,377.00 (47.6%). Check: 359,409 + 189,929 + 499,377 = 1,048,715 ✓.

Priority callout: $279,990 of the $359,409 high-risk ARR (C-0B7D2C30, C-0BCDB8C2, C-0D2AB865, C-0BBE3E60, C-0F5D2323) is all multi-year-contract accounts with double-digit usage declines renewing within the next 5 days to 5 weeks.
ticket-theme-synthesis0.80
Plain-English question: What are the real themes in this quarter's support tickets, ignoring the existing tags?
Prompt — what the model saw
Synthesize the quarter's support tickets into 4 to 7 themes. Do not trust existing tags; read the text. Per theme: name, count, share, distinct accounts, ARR affected, two ticket ids, one-line recommendation. Separate broad patterns from single-account noise and rank by ARR exposure, not volume.
cost $0.0959 · 240s · in 2,932 / out 12,003 tok
✗ theme_count_range✓ counts_verifiable✓ single_account_flagged✓ arr_ranked✓ ticket_ids

✓ Correct Answer (ground truth)

Plain-English answer: 80 tickets in 4-7 themes — points-not-posting leads (20 tickets, 25%, $31,100 ARR affected) — each with counts, distinct accounts, ARR, and a recommendation.
The answer in one line: total_tickets = 80; quarter_window = 2026-06-01 through 2026-08-31
total_tickets80
noteexisting tags are noisy/unreliable; classify by text
quarter_window2026-06-01 through 2026-08-31
ranking_by_arr_exposurehris_provisioning, redemption_failures, billing_dispute, points_not_posting, slack_sync
sample_ticket_ids80 items — e.g. IC-460004, IC-460016, IC-460001, IC-460041
single_account_noise_themebilling_dispute
themes{"points_not_posting": {"count": 20, "share_pct": 25.0, "distinct_accounts": 9, "arr_affected": "31100.00", "example_ticket_ids": ["IC-460001", "IC-460002"], "single_account": false}, "slack_sync": {"
graded checkstheme_count_range · counts_verifiable · single_account_flagged · arr_ranked · ticket_ids

Model's actual answer

show the model's full answer
**Method note:** 80 tickets total, 24 distinct accounts, $285,400 total ARR represented. "ARR affected" = sum of ARR across *distinct* accounts in the theme (per-ticket ARR is not summed, to avoid double-counting repeat submitters). Tags ignored — themes built from body text only.

**Ranked by ARR exposure:**

**1. HRIS provisioning/sync failures — BROAD PATTERN**
- Count: 12 tickets (15.0% of quarter)
- Distinct accounts: 3 — C-0B2213A9 ($36,000), C-0DDFC9A7 ($48,000), C-0F6C0F34 ($30,000)
- ARR affected: 36,000 + 48,000 + 30,000 = **$114,000** (39.9% of represented ARR)
- Ticket IDs: IC-460059, IC-460062
- Recommendation: Escalate to engineering as P1 — silent sync failures ("log shows no errors") blocking new-hire onboarding at your three largest accounts in this set.

**2. Redemption/checkout failures (gift cards) — BROAD PATTERN**
- Count: 18 tickets (22.5%)
- Distinct accounts: 7 — C-0CEF69FD ($8,900), C-0B827671 ($10,700), C-0FCCD2DF ($9,900), C-0F876796 ($8,700), C-14264ABD ($11,000), C-0D9CA315 ($9,900), C-0B0F1BAB ($10,300)
- ARR affected: **$69,400**
- Ticket IDs: IC-460025, IC-460024
- Recommendation: Audit the redemption pipeline for the points-deducted-but-no-card case — it is the only theme where customers lose value already paid for, which is a refund/compliance exposure, not just a bug.

**3. Recurring invoice/seat-count billing errors — SINGLE-ACCOUNT NOISE**
- Count: 16 tickets (20.0%)
- Distinct accounts: 1 — C-0E9C27D1 ($52,000)
- ARR affected: **$52,000**
- Ticket IDs: IC-460071, IC-460069
- Recommendation: Not a product pattern — this is one $52K account filing the same seat-count complaint (200 seats billed vs. 150 licensed) roughly every 3 days for a full quarter; route to billing escalation/CSM immediately as a churn risk, not to support triage.

**4. Points not posting / missing balances — BROAD PATTERN**
- Count: 20 tickets (25.0% — highest volume of any theme)
- Distinct accounts: 9 — C-0D0B047C ($4,500), C-0BF20542 ($4,500), C-0D3278C7 ($3,500), C-0D6CC8E3 ($4,200), C-0D284E42 ($3,400), C-0BE96399 ($2,700), C-21FEBCBB ($2,900), C-0DD0626C ($2,500), C-0B2895EF ($2,900)
- ARR affected: **$31,100**
- Ticket IDs: IC-460004, IC-460001
- Recommendation: Highest ticket volume but lowest-value accounts — investigate the recognition-to-balance posting job (multiple "delivered but never arrived" reports suggest an async processing failure), and check whether weekend batch runs are the trigger given repeated "after the weekend" phrasing.

**5. Slack integration failures — BROAD PATTERN**
- Count: 14 tickets (17.5%)
- Distinct accounts: 4 — C-10A56B0F ($5,400), C-8C2E8F00 ($5,200), C-0B843542 ($4,400), C-0BA71F12 ($3,900)
- ARR affected: **$18,900**
- Ticket IDs: IC-460047, IC-460046
- Recommendation: Three distinct symptom clusters (sync toggle self-resetting, re-auth not sticking, slash-command errors) across 4 accounts point to an OAuth/token-refresh defect — likely one root cause, so fix that before patching symptoms individually.

**Arithmetic check:** 12 + 18 + 16 + 20 + 14 = 80 tickets ✓; $114,000 + $69,400 + $52,000 + $31,100 + $18,900 = $285,400 ✓

**Bottom line:** HRIS provisioning is the top ARR exposure despite being the *lowest* ticket volume of the five themes — volume-based triage would have deprioritized it. The billing cluster (theme 3) is the only single-account issue and should be handled as an account escalation, not a support queue.
marketing 4 tests
case-study-match0.75
Plain-English question: Which three existing customers are the best case-study matches for this prospect?
Prompt — what the model saw
Given the prospect, rank the three most similar existing customers with public case studies for use as social proof. Similarity on industry, size band, use case, region. Cite the fields that drove each match. Do not name any customer without a case study and never include billing data or contact names.
cost $0.0318 · 60s · in 0 / out 2,837 tok
✓ top3✗ decoys_excluded✓ cites_fields✓ no_billing

✓ Correct Answer (ground truth)

Plain-English answer: Three ranked matches on industry, size, use case, and region (e.g. C-11C31562), with the higher-scoring decoy explicitly excluded for a mismatched field.
The answer in one line: prospect_alias = C-82AF3719
constraintsno_billing_data, no_contact_names, only_customers_with_has_case_study_true_named
decoys_excludedcompany_alias=C-D6217CAA; similarity_score=4; reason=no_case_study | company_alias=C-C153868F; similarity_score=4; reason=no_case_study
prospect_aliasC-82AF3719
prospect_fields_usedindustry, size_band, use_case, region
top3company_alias=C-11C31562; similarity_score=3 | company_alias=C-64171065; similarity_score=3 | company_alias=C-A13C193D; similarity_score=3
graded checkstop3 · decoys_excluded · cites_fields · no_billing

Model's actual answer

show the model's full answer
**Prospect: C-82AF3719 — Technology / Mid-Market / employee_recognition / NA-West**

**Method:** 1 point per matching dimension (industry, size_band, use_case, region), max 4. Only customers with has_case_study=true are eligible — 8 of 25 qualify; 17 excluded.

**Scoring (all 8 case-study customers):**

| Customer | Industry | Size | Use case | Region | Score |
|---|---|---|---|---|---|
| C-64171065 | Tech ✓ | MM ✓ | emp_recognition ✓ | NA-East ✗ | 3/4 |
| C-11C31562 | Mfg ✗ | MM ✓ | emp_recognition ✓ | NA-West ✓ | 3/4 |
| C-A13C193D | Tech ✓ | MM ✓ | retention ✗ | NA-West ✓ | 3/4 |
| C-CD4829A7 | Tech ✓ | MM ✓ | retention ✗ | NA-East ✗ | 2/4 |
| C-9D69D147 | FinServ ✗ | MM ✓ | retention ✗ | NA-West ✓ | 2/4 |
| C-FEFCC020 | ✗ | ✗ | ✗ | ✗ | 0/4 |
| C-11209339 | ✗ | ✗ | ✗ | ✗ | 0/4 |
| C-0CE315B2 | ✗ | ✗ | ✗ | ✗ | 0/4 |

**Tiebreak among the three 3/4 matches:** weight use_case and industry above region for social-proof relevance. Missing region is least damaging; missing use_case is most damaging for a use-case-specific purchase.

**Ranking:**

1. **C-64171065** — matches industry (Technology), size_band (Mid-Market), use_case (employee_recognition); misses only region (NA-East vs NA-West).
2. **C-11C31562** — matches use_case (employee_recognition), size_band (Mid-Market), region (NA-West); misses industry (Manufacturing vs Technology).
3. **C-A13C193D** — matches industry (Technology), size_band (Mid-Market), region (NA-West); misses use_case (retention vs employee_recognition).

**Data limitations (explicit):**
- No case-study customer matches all 4 dimensions. Two perfect 4/4 matches exist (C-C153868F, C-D6217CAA — Technology/Mid-Market/employee_recognition/NA-West) but both have has_case_study=false and are therefore excluded per your rule.
- The dataset contains no billing data or contact names, so none are included. No other customer attributes (e.g., headcount, ARR, tenure) were available to break the 3/3/3 tie beyond the stated weighting.
channel-attribution-mathPASS 1.00
Plain-English question: Compute cost per SQM, cost per SQO, and pipeline per dollar for every channel over six months.
Prompt — what the model saw
Compute channel performance for the trailing 6 months. Paid: spend, SQMs, SQOs, cost per SQM, cost per SQO, SQM-to-SQO rate, pipeline amount, pipeline per dollar. Organic: volume, SQO rate, pipeline. Report a channel with spend and zero SQMs as undefined, not zero. Flag rows where SQO date precedes SQM date. Finish with a reallocation recommendation and your confidence given sample sizes.
cost $0.1030 · 183s · in 6,272 / out 8,818 tok
✓ undefined_channel✓ undefined_not_zero✓ flagged_rows✓ best_channel✓ reallocation

✓ Correct Answer (ground truth)

Plain-English answer: LinkedIn Ads at $960/SQM and $3,000/SQO, organic search at $90,000 pipeline on zero spend, paid social flagged at $18,000 with zero SQMs, and a channel winner named.
The answer in one line: best_pipeline_per_dollar_channel = paid_search
best_pipeline_per_dollar_channelpaid_search
confidence_notewebinars and referral have small SQM samples (<20); treat rates as directional.
flagged_rows_sqo_before_sqmCT-000041, CT-000044
months2026-03, 2026-04, 2026-05, 2026-06, 2026-07, 2026-08
per_channel{"linkedin_ads": {"spend_usd": 24000, "sqms": 25, "sqos": 8, "pipeline_amount": 96000, "cost_per_sqm": 960.0, "cost_per_sqo": 3000.0, "sqm_to_sqo_rate": 0.32, "pipeline_per_dollar": 4.0}, "organic_sea
reallocation_recommendationfrom=paid_social; to=paid_search; rationale=paid_social has spend but zero SQMs (cost per SQM undefined, not zero); paid_search has the highest pipeline per dollar.
undefined_channelspaid_social
graded checksundefined_channel · undefined_not_zero · flagged_rows · best_channel · reallocation

Model's actual answer

show the model's full answer
TRAILING 6 MONTHS (2026-03 through 2026-08)

PAID CHANNELS

paid_search
- Spend: 6 × $6,000 = $36,000
- SQMs: 40 | SQOs: 18
- Cost per SQM: 36,000 / 40 = $900.00
- Cost per SQO: 36,000 / 18 = $2,000.00
- SQM→SQO rate: 18 / 40 = 45.0%
- Pipeline: 18 × $40,000 = $720,000
- Pipeline per dollar: 720,000 / 36,000 = $20.00

linkedin_ads
- Spend: 6 × $4,000 = $24,000
- SQMs: 25 | SQOs: 8
- Cost per SQM: 24,000 / 25 = $960.00
- Cost per SQO: 24,000 / 8 = $3,000.00
- SQM→SQO rate: 8 / 25 = 32.0%
- Pipeline: 8 × $12,000 = $96,000
- Pipeline per dollar: 96,000 / 24,000 = $4.00

paid_social
- Spend: 6 × $3,000 = $18,000
- SQMs: 0 | SQOs: 0 | Pipeline: $0
- Cost per SQM, cost per SQO, conversion rate, pipeline per dollar: UNDEFINED (spend with zero SQMs — reported as undefined, not zero). Either attribution/tracking is broken for this channel or it produced nothing; the data cannot distinguish the two.

webinars
- Spend: 6 × $1,500 = $9,000
- SQMs: 12 | SQOs: 5
- Cost per SQM: 9,000 / 12 = $750.00
- Cost per SQO: 9,000 / 5 = $1,800.00
- SQM→SQO rate: 5 / 12 = 41.7%
- Pipeline: 5 × $12,000 = $60,000
- Pipeline per dollar: 60,000 / 9,000 = $6.67

ORGANIC CHANNELS (no spend tracked in channel_spend.csv)

organic_search
- Volume: 30 SQMs | SQO rate: 10 / 30 = 33.3% | Pipeline: 10 × $9,000 = $90,000

referral
- Volume: 15 SQMs | SQO rate: 6 / 15 = 40.0% | Pipeline: 6 × $8,000 = $48,000

DATA-QUALITY FLAGS (SQO date precedes SQM date — impossible sequence, likely backdated or mis-attributed):
- CT-000044, linkedin_ads: SQM 2026-07-23, SQO 2026-07-18
- CT-000041, linkedin_ads: SQM 2026-06-14, SQO 2026-06-09
Both rows are in linkedin_ads; they inflate that channel's SQO count and pipeline. Excluding them would drop linkedin_ads to 6 SQOs, $72,000 pipeline, $4,000/SQO, $3.00 pipe/$.

REALLOCATION RECOMMENDATION
1. Hold or modestly increase paid_search — best pipeline efficiency among paid channels ($20 pipe/$, 45% SQM→SQO, $2,000/SQO) on the largest sample.
2. Cut or pause paid_social's $18,000 pending an attribution audit — six months of spend with zero tracked SQMs cannot be justified on this data. Do not treat it as "zero performance" until tracking is verified; if tracking is confirmed intact, reallocate the full $18,000.
3. Reallocate a portion of linkedin_ads budget to webinars — webinars beats it on every paid metric ($1,800 vs $3,000 per SQO; $6.67 vs $4.00 pipe/$), and linkedin's two flagged rows further weaken its case.
4. Keep organic and referral as-is — no spend data, so no reallocation basis.

CONFIDENCE
- paid_search: moderate-high (40 SQMs, 18 SQOs).
- webinars and linkedin_ads: low-moderate (12/5 and 25/8 — small SQO counts; one flagged row each in linkedin is 12.5% of its SQOs).
- paid_social: no basis to judge (zero tracked outcomes).
- Organic channels: low-moderate (30 and 15 SQMs, but no spend data to compute efficiency).
Overall: moderate. The direction (search strong, linkedin weak, paid_social unjustifiable as-is) is well supported; the precise split of any reallocated dollars is not, given the small SQO counts.
competitor-battlecard-updatePASS 1.00
Plain-English question: Update the Rivally battlecard with sourced pricing and where each side wins.
Prompt — what the model saw
Update the battlecard for the competitor. Sections: one-line positioning, pricing with source and date (newer source wins, note the conflict), where they win, where we win, objections and responses, recent changes, our 12-month win/loss record against them. Cite a snippet id for every factual claim. Rep opinion on a call is not a fact about the competitor. Anything from the old card you cannot re-source gets marked unverified.
cost $0.0464 · 102s · in 310 / out 5,063 tok
✓ competitor_named✓ pricing✓ win_loss✓ unverified_marked✓ citations

✓ Correct Answer (ground truth)

Plain-English answer: Rivally at $7 per user per month (the August source wins over the stale $5, conflict noted), a 13-7 win/loss record, and two unverified old claims flagged.
The answer in one line: competitor = Rivally
citation_ids_available25 items — e.g. S01, S02, S03, S04
competitorRivally
pricingvalue_per_user_per_month=7; source_snippet=S17; source_date=2026-08-12
rep_opinion_snippet_ids_not_factsS09, S21
unverified_old_claimsRivally lacks a Slack integration., Rivally was acquired by WorkHuman in 2025.
win_loss_12mowins=13; losses=7
graded checkscompetitor_named · pricing · win_loss · unverified_marked · citations

Model's actual answer

show the model's full answer
BATTLECARD: RIVALLY — updated from snippets S01–S25 (latest data: 2026-09-03). All claims cite snippet ids. Old-card items that could not be re-sourced are flagged in Section 8.

1) ONE-LINE POSITIONING
Points-based recognition platform for mid-market and distributed EU teams, now expanding into enterprise/EU (Dublin office, EU data residency GA — S15; ex-Workday VP EMEA hire — S11) and into engagement surveys via the "Rivally Pulse" add-on (S06, S23). Points-based feed confirmed by S02; mid-market presence by S04.

2) PRICING (newer source wins; conflict noted)
- CURRENT: $7 per user/month, Recognition Starter tier, annual billing required. Source: pricing page, 2026-08-12 (S17). Corroborated by deal mention 2026-08-14: $7/user/mo list (S18).
- CONFLICT / price history:
  - $5/user/mo, annual billing — pricing page 2026-01-20 (S03); still $5 on 2026-04-01 (S08).
  - $6.50/user/mo quoted to a 500-seat prospect, annual term — deal mention 2026-06-02 (S13).
  - $7/user/mo — pricing page 2026-08-12 (S17). Newest pricing-page source wins over S03/S08 and over the old card's "$5 as of 2026-01."
- Discounting: only confirmed data point is 15% off list for a 3-year term, per deal mention 2026-08-14 (S18). The claim that Rivally is "discounting aggressively" is rep opinion, not confirmed (S21) — excluded as a fact.
- Add-on: Rivally Pulse is priced separately, not bundled (S23, 2026-09-01). No price for Pulse is given in any snippet.

3) WHERE THEY WIN
- EU / distributed teams: multi-language support praised by an EU enterprise reviewer (S12); EU data residency pitched to a prospect (S05) and generally available (S15); Dublin office (S15); dedicated EMEA leadership hire (S11).
- Fast time-to-value: mid-market reviewer reports setup under a week with Slack integration working out of the box (S04).
- Engagement: recognition feed praised as engaging (S02, S16).
- Support: response time under 4 hours praised (S22).
- Ecosystem: Microsoft Teams app v2 in public preview (S19).

4) WHERE WE WIN
- Analytics depth: reviewers note limited analytics (S02), basic dashboards vs. enterprise tools (S07), CSV-only analytics exports (S20). An 800-seat prospect picked Bonusly over Rivally citing analytics depth (S25).
- Enterprise admin tooling: no SCIM provisioning, manual user management painful (S10); admin tooling lags peers (S16); no bulk recognition editing (S24).
- EMEA rewards: rewards catalog in EMEA thinner than the US catalog (S14).
- Lock-in friction: CSV-only analytics exports made migration off Rivally hard for one reviewer (S20).
- Pricing trajectory: list price rose from $5 (S03, 2026-01) to $7 (S17, 2026-08) in ~7 months; Pulse is a paid add-on on top (S23).

5) OBJECTIONS AND RESPONSES
- Objection: "Rivally offers EU data residency." (Prospect-facing: S05; GA: S15.)
  Response: We have no sourced data on our own EU data residency in the provided snippets — no counter available from this data. Adjacent sourced point: their EMEA rewards catalog is thinner than US (S14).
- Objection: "Rivally quoted us less / will discount." (Confirmed data: 15% off for a 3-year term, S18.)
  Response: Their list price went $5 → $6.50 quoted → $7 within seven months (S03, S13, S17); the only confirmed discount is tied to a 3-year commitment (S18). Note: "aggressive discounting" is rep opinion (S21) and cannot be used as a fact.
- Objection: "Rivally sets up in under a week and Slack works out of the box." (S04.)
  Response: No sourced data on our own setup time or Slack integration in the provided snippets — no counter available from this data.
- Objection: "Their support responds in under 4 hours." (S22.)
  Response: No sourced data on our support response times — no counter available from this data.
- Objection: "Their feed is more engaging." (S02, S16.)
  Response: Counter with admin/enterprise gaps: no SCIM (S10), no bulk recognition editing (S24), basic dashboards (S07), CSV-only exports (S20).

6) RECENT CHANGES (Jun–Sep 2026)
- 2026-06-02: quoted $6.50/user/mo to a 500-seat prospect (S13).
- 2026-07-01: Dublin office opened; EU data residency GA (S15).
- 2026-08-12: price page updated to $7/user/mo (S17); 2026-08-14: 15% discount for 3-year term observed (S18).
- 2026-08-20: Microsoft Teams app v2 public preview (S19).
- 2026-09-01: Rivally Pulse exits beta, sold as a paid add-on (S23); originally launched 2026-03-05 (S06).
- Context: $40M Series C led by Northgate Ventures, 2025-11-04 (S01).

7) 12-MONTH WIN/LOSS RECORD VS. RIVALLY (deals_with_competitor.csv, 2025-09 through 2026-08)
- Wins (13): Deal-072E31 (2025-09), Deal-A9FD43 (2025-10), Deal-F65C8F (2025-10), Deal-7AA785 (2025-11), Deal-44C524 (2025-12), Deal-0D0CD6 (2026-01), Deal-E46EAB (2026-01), Deal-D5B790 (2026-02), Deal-1D2392 (2026-02), Deal-5C636E (2026-03), Deal-67BE14 (2026-06), Deal-1B6969 (2026-07), Deal-F03E7B (2026-08).
- Losses (7): Deal-7767F5 (2025-09), Deal-D263E0 (2025-11), Deal-935746 (2025-12), Deal-9066A6 (2026-03), Deal-5645A5 (2026-04), Deal-72A02F (2026-04), Deal-C6FFAA (2026-05).
- Arithmetic: 13 wins + 7 losses = 20 deals; 13 / 20 = 0.65 → 65% win rate.
- Trend: worst stretch was 2026-04 through 2026-05 (0 wins, 3 losses), followed by 3 consecutive wins 2026-06 through 2026-08. The deals file gives no loss reasons; the only sourced win reason is the 800-seat analytics-depth win (S25, 2026-09-03), which is a snippet, not a row in the deals file, and is therefore not counted in the 13–7 record.

8) UNVERIFIED / EXCLUDED FROM OLD CARD
- "Rivally was acquired by WorkHuman in 2025" — UNVERIFIED. No snippet supports an acquisition. The only ownership-related fact is the $40M Series C led by Northgate Ventures (S01), which does not confirm or refute an acquisition.
- "Rivally lacks a Slack integration" — CONTRADICTED and removed. A G2 reviewer reports the Slack integration worked out of the box (S04, 2026-02-02).
- "Pricing starts at $5/user/mo" — outdated, superseded by S17 ($7, 2026-08-12).
- "Strong in EU enterprise with multi-language support" — retained, re-sourced to S12 (plus S05, S11, S15).
- Excluded as non-facts (rep opinion on calls): "UI is clunky" (S09, explicitly not verified against the product); "discounting aggressively" (S21, not confirmed).

DATA GAPS (explicit): No snippets provide Bonusly's own pricing, setup time, support SLAs, or EU data residency status, so no sourced counters exist for those objections. No loss reasons exist in the deals file. No snippet gives Rivally Pulse's price, enterprise-tier pricing, or free-trial terms.
weekly-marketing-goals-rollupPASS 1.00
Plain-English question: Are we on pace on the quarter's marketing goals?
Prompt — what the model saw
Produce the weekly marketing goals update for the active quarter: SQMs, SQOs, DS2s, closed-lost MIA rate, same-quarter close count, active pipeline coverage against target. For each: QTD actual, target, delta, pace (on/behind/ahead given days elapsed). One paragraph on what moved this week.
cost $0.0318 · 58s · in 0 / out 2,873 tok
✓ metrics_present✓ deltas✓ coverage_ratio✓ pace✓ paragraph

✓ Correct Answer (ground truth)

Plain-English answer: SQMs ahead of pace, SQOs on pace, DS2s behind (35 short of expected at day 66 of 92), with deltas and same-quarter closes.
The answer in one line: days_in_quarter = 92; days_elapsed = 66; coverage_ratio_active_pipeline = 0.75; elapsed_fraction = 0.7174 (+1 more below)
days_in_quarter92
coverage_ratio_active_pipeline0.75
days_elapsed66
elapsed_fraction0.7174
metrics{"SQMs": {"qtd_actual": 230, "target": 300, "delta": -70, "expected_at_day_elapsed": 215.22, "pace": "ahead"}, "SQOs": {"qtd_actual": 84, "target": 120, "delta": -36, "expected_at_day_elapsed": 86.09,
pace_ruleexpected = target * days_elapsed/days_in_quarter; within +/-5% of expected = on, above = ahead, below = behind; lower_better metrics invert.
this_week_paragraph_required1
graded checksmetrics_present · deltas · coverage_ratio · pace · paragraph

Model's actual answer

show the model's full answer
WEEKLY MARKETING GOALS UPDATE — Q3-2026 (day 66 of 92; 66/92 = 71.7% elapsed; 26 days remaining)

Pro-rata expectation at day 66 = target × 66/92. Pace is judged vs. that bar.

METRICS
--------------------------------------------------------------------
SQMs
  QTD actual: 230 | Target: 300 | Delta: 230 − 300 = −70 (76.7% of target)
  Pro-rata: 300 × 66/92 = 215.2 → 230 vs 215.2 = +14.8 → AHEAD

SQOs
  QTD actual: 84 | Target: 120 | Delta: 84 − 120 = −36 (70.0% of target)
  Pro-rata: 120 × 66/92 = 86.1 → 84 vs 86.1 = −2.1 → BEHIND (marginal)

DS2s
  QTD actual: 40 | Target: 75 | Delta: 40 − 75 = −35 (53.3% of target)
  Pro-rata: 75 × 66/92 = 53.8 → 40 vs 53.8 = −13.8 → BEHIND

Closed-lost MIA rate
  QTD actual: 5 / 25 = 0.20 (20%) | Target: 0.10 (10%), lower-better
  Delta: +10pp (2× target) → BEHIND (a rate already at 2× target with 26 days left; pro-rata does not apply to lower-better rates — any reading above 10% is off-track)

Same-quarter closes
  QTD actual: 10 | Target: 20 | Delta: 10 − 20 = −10 (50.0% of target)
  Pro-rata: 20 × 66/92 = 14.3 → 10 vs 14.3 = −4.3 → BEHIND

Active pipeline coverage
  QTD actual: $3,000,000 | Target: $4,000,000 | Delta: −$1,000,000 (coverage = 3.0M / 4.0M = 0.75×)
  Pro-rata: 4.0M × 66/92 = $2.87M → nominally ahead by $130k, but pipeline is a stock/coverage metric where the bar is the full $4M → BEHIND on coverage basis (0.75×)

WHAT MOVED THIS WEEK
Week-over-week movement cannot be computed from the data provided — the files contain only a single QTD snapshot with no prior-week figures, so no weekly delta is available and I won't invent one. On a quarter-to-date basis, the standing picture is: top of funnel is the only metric ahead of pace (SQMs +14.8 vs pro-rata), while everything downstream lags — SQOs are marginally behind (−2.1), DS2s materially behind (−13.8, 53.3% of target), same-quarter closes at exactly half of target (10 vs 20), and the closed-lost MIA rate sits at 20%, double the 10% goal (5 of 25 closed-lost). QTD conversion arithmetic: SQM→SQO = 84/230 = 36.5%; SQO→DS2 = 40/84 = 47.6%. Pipeline coverage is 0.75× against the $4M target. To report actual weekly movement next cycle, a prior-week snapshot (or week-stamped history) of the same metrics is needed.
communication 5 tests
ceo-slack-compressionPASS 1.00
Plain-English question: Compress the forecast analysis into one Slack message the CEO can read in ten seconds.
Prompt — what the model saw
Compress the forecast analysis below into a single Slack message for the CEO. Rules: at most three sentences; under 80 words; lead with the decision or the so-what; keep every distinct number that matters and the one caveat that changes the conclusion; cut any sentence that does not carry a claim, a number, or a directive. No em dashes, no throat-clearing phrases.
cost $0.0184 · 19s · in 0 / out 602 tok
✓ max_3_sentences✓ under_80_words✓ weighted_present✓ caveat_present✓ no_em_dash

✓ Correct Answer (ground truth)

Plain-English answer: Three sentences, under 80 words, decision first, every distinct number kept plus the one caveat — all checked by deterministic text rules.
The answer in one line: weighted_number = 115976; must_keep = data-quality caveat — deleting the bad news fails
weighted_number115,976
limitmax 3 sentences, max 80 words, no em-dashes
must_keepdata-quality caveat — deleting the bad news fails
graded checksmax_3_sentences · under_80_words · weighted_present · caveat_present · no_em_dash

Model's actual answer

show the model's full answer
Q3 weighted forecast is 115,977: 44,729 COMMIT + 71,248 (35% of 203,565 BEST_CASE), with PIPELINE counted at zero. 32 deals worth 227,575 sit outside the quarter, including COMMIT Deal-D348E1 at 13,770. Caveat: owner is blank on 85 of 86 deals and why-buys is missing on all 7 COMMIT deals, so this number looks precise but can't be validated by rep or rationale.
stalled-deal-followup-emailPASS 1.00
Plain-English question: Write a follow-up email for this stalled deal — under 120 words, one prior touch, one ask.
Prompt — what the model saw
Write a follow-up email for this stalled deal. Under 120 words. Reference one specific prior touch from the engagement history, make one clear ask, and match the tone of the prior thread. No invented details.
cost $0.0397 · 82s · in 0 / out 4,182 tok
✓ word_cap✓ references_touch✓ single_ask

✓ Correct Answer (ground truth)

Plain-English answer: An email that references the July 28 milestone-awards demo, makes one clear ask, and invents nothing — Deal-0D2F7A, 31 days since last engagement.
The answer in one line: amount = 5100; word_cap = 120; days_since_last_engagement = 31; deal_alias = Deal-0D2F7A (+4 more below)
amount5,100
days_since_last_engagement31
word_cap120
allowed_facts14 items — e.g. Deal-0D2F7A, $5,100, DS3, Alex Franklin
allowed_newENG-00001, ENG-00002, ENG-00003, ENG-00004, ENG-00005
deal_aliasDeal-0D2F7A
must_referenceengagement_id=ENG-00002; date=2026-07-28; type=meeting; why=most substantive prior touch; concrete prospect reactions to milestone awards and points catalog
ownerAlex Franklin
single_cta1
stageDS3
tonewarm, concise, professional (matches prior thread)
graded checksword_cap · references_touch · single_ask

Model's actual answer

show the model's full answer
Subject: Following up on pricing for 150 seats (Deal-0D2F7A)

Hi [Name],

I'm following up on the pricing recap I sent on August 5 for the 150 seats we discussed in your demo. I want to make sure the pricing question your team raised gets a clear answer rather than sitting in an inbox.

Would you have 20 minutes in the next week or two to walk through the recap together? Happy to work around your calendar.

Best,
Alex Franklin

---
Notes:
- Word count: 70 body words (78 with subject line) — under 120.
- Prior touch referenced: ENG-00001 (2026-08-05 email, pricing recap for 150 seats, no reply).
- One ask: a 20-minute call to review the pricing recap.
- Tone matched to prior outbound emails: concise, professional, recap-style.
- Recipient name is not in the provided data — placeholder [Name] used rather than inventing one.
funnel-digest-paragraphPASS 1.00
Plain-English question: Write the weekly GTM digest — one paragraph per team, every number traceable.
Prompt — what the model saw
Write the weekly GTM digest: one paragraph per team (Marketing, Sales, CS, PLG) in Bonusly brand voice. Every number must trace to the extracts. No customer contact names, no employee performance commentary, no compensation. Four paragraphs, under 400 words total.
cost $0.0398 · 40s · in 0 / out 1,790 tok
✓ four_paragraphs✓ word_cap✓ teams

✓ Correct Answer (ground truth)

Plain-English answer: Four paragraphs (Marketing, Sales, CS, PLG) using only the allowed numbers, no contact names, no performance commentary, under 400 words.
The answer in one line: word_cap = 400
word_cap400
allowed_numbers16 items — e.g. 2, 3, 9, 08
bannedcustomer contact names, employee performance commentary, compensation
required_paragraphsMarketing, Sales, CS, PLG
graded checksfour_paragraphs · word_cap · teams

Model's actual answer

show the model's full answer
Note: the bonusly-brand skill failed to load (tool error), so voice is applied from known Bonusly conventions.

---

**Bonusly Weekly GTM Digest — Week of 2026-08-31**

**Marketing:** Marketing wrapped the week at 46 SQMs against a 52 target — 88% attainment (46 ÷ 52 = 0.885), a shortfall of 6. The webinar channel was the bright spot, delivering 18 of those SQMs, or 39% of the weekly total (18 ÷ 46 = 0.391). Paid search held its efficiency at $150 cost per SQM, keeping acquisition economics steady even as volume came in light. A focused push on the webinar motion looks like the clearest lever to close the gap next week.

**Sales:** Sales converted 14 SQOs and booked 9 DS2 meetings for the week. New pipeline created totaled $310,000 — about $22.1K of pipeline per SQO ($310,000 ÷ 14 ≈ $22,143). Three same-quarter deals closed, a solid contribution to in-quarter momentum. With 9 DS2 meetings on the calendar, the team heads into next week with healthy mid-funnel coverage.

**CS:** CS saved 2 renewals this week — two relationships kept strong through focused attention. Team NPS moved to 61, a signal that customers are feeling the value. Three red-flag accounts remain open heading into next week, and they represent the priority watch list; proactive outreach on those accounts is the top CS motion for the coming days.

**PLG:** PLG welcomed 412 new signups this week, with activation holding at 31% — roughly 128 activated users (412 × 0.31 ≈ 128). Thirty-eight companies reached the aha moment of 10 recognition gives, crossing the threshold where recognition starts to stick. Growing that aha-moment cohort is the natural next milestone for the product-led motion.

---

Word count: ~290. All figures trace directly to the extracts; derived figures (88% attainment, 39% webinar share, $22.1K per SQO, ~128 activated) are shown with their arithmetic. No data was available on pipeline coverage, SQO conversion rate vs. prior weeks, or churn figures — those metrics weren't in the extracts, so they're not cited.
partner-digest-sectionPASS 1.00
Plain-English question: Write the partner digest so every partner in the roster appears exactly once.
Prompt — what the model saw
Write the partner digest for the period. Every partner in the roster appears once: active plays get status with numbers, quiet partners get one line. Partner-sourced pipeline count and amount must match the deal data. Nothing omitted, nothing added.
cost $0.0237 · 34s · in 0 / out 1,511 tok
✓ all_partners✓ pipeline_amount✓ quiet_one_line

✓ Correct Answer (ground truth)

Plain-English answer: Four active partners with status and numbers, four quiet partners with one line each, and partner-sourced pipeline that reconciles to the deal data.
The answer in one line: rule = every roster partner named exactly once; quiet partners get one line; no partners outside roster.
active_no_dealsWorkWell Group
partner_sourced_pipelinecount=5; amount=550000
per_partner_pipeline{"Apex Rewards Co": {"count": 2, "amount": 275000}, "HRCloud Partners": {"count": 1, "amount": 140000}, "CultureBridge": {"count": 2, "amount": 135000}}
quiet_partnersRecogniq, KudosWave, PeopleFirst Advisors, TotalPerk
rosterpartner_name=Apex Rewards Co; status=active | partner_name=HRCloud Partners; status=active | partner_name=CultureBridge; status=active | partner_name=WorkWell Group; status=active | partner_name=Recogniq; status=quiet | partner_name=KudosWave; status=quiet … (+2 more)
ruleevery roster partner named exactly once; quiet partners get one line; no partners outside roster.
graded checksall_partners · pipeline_amount · quiet_one_line

Model's actual answer

show the model's full answer
PARTNER DIGEST — Period: 2026-08-24 to 2026-09-04

ACTIVE PARTNERS

Apex Rewards Co
- Partner-sourced pipeline: 2 deals, $275,000 (Deal-DDAAF2 $180,000 + Deal-2CF33E $95,000 = $275,000)
- Status: Two sourced opps this period, both in DS1; confirmed logged with UTM Source = Partner (Slack, 2026-08-25)
- Co-webinar locked for 09-15

HRCloud Partners
- Partner-sourced pipeline: 1 deal, $140,000 (Deal-F1CDA5)
- Status: One sourced opp this period; integration referral closed security review and moved to DS2 (Slack, 2026-08-27)

CultureBridge
- Partner-sourced pipeline: 2 deals, $135,000 (Deal-067213 $60,000 + Deal-096E1D $75,000 = $135,000)
- Status: Two sourced opps from the lunch-and-learn, both early stage (Slack, 2026-08-29)

WorkWell Group
- Partner-sourced pipeline: 0 deals, $0
- Status: No sourced deals this period; joint playbook restart targeted for Q4, planning call booked for 09-09 (Slack, 2026-09-02)

QUIET PARTNERS

Recogniq — No activity this period.
KudosWave — No activity this period.
PeopleFirst Advisors — No activity this period.
TotalPerk — No activity this period.

TOTALS

Partner-sourced deals: 5 (Apex 2 + HRCloud Partners 1 + CultureBridge 2)
Partner-sourced pipeline: $550,000 ($275,000 + $140,000 + $135,000)

DATA NOTES
- Deal counts and amounts are taken from partner_deals.csv; they reconcile with the Slack export (Apex 2, HRCloud 1, CultureBridge 2, WorkWell 0).
- Deal stages (DS1, DS2, early stage) are stated only in the Slack export; partner_deals.csv contains no stage or date fields, so stage attribution relies on Slack.
- All 8 roster partners appear once; quiet partners have no deals in the deal file and no messages in the export.
rfp-answer-from-policyPASS 1.00
Plain-English question: Answer the 10 RFP items from the policy excerpts only — no guessing.
Prompt — what the model saw
Answer the 10 questionnaire items from the policy excerpts only. Cite the excerpt id for every answer. Where the excerpts do not answer the question, mark it unanswerable and say what document would; do not guess. Tone: vendor response, direct.
cost $0.0195 · 21s · in 0 / out 812 tok
✓ answered_count✓ unanswerable_flagged✓ citations✓ no_guess

✓ Correct Answer (ground truth)

Plain-English answer: Seven answered with excerpt citations (Q1-Q7), three flagged unanswerable with the document that would answer them (Q8-Q10).
The answer in one line: answered_count = 7; rule = unanswerable items must be flagged and the document that would answer them named; do not guess.
answered_count7
answer_map{"Q1": ["P1"], "Q2": ["P1"], "Q3": ["P2"], "Q4": ["P3"], "Q5": ["P4"], "Q6": ["P5"], "Q7": ["P6"]}
ruleunanswerable items must be flagged and the document that would answer them named; do not guess.
unanswerableQ8, Q9, Q10
valid_citation_idsP1, P2, P3, P4, P5, P6, P7, P8
graded checksanswered_count · unanswerable_flagged · citations · no_guess

Model's actual answer

show the model's full answer
Q1. Yes. Customer data is encrypted at rest with AES-256. [P1]

Q2. Yes. Data in transit is encrypted with TLS 1.2 or higher. [P1]

Q3. Yes. Bonusly maintains a SOC 2 Type II report, renewed annually. A copy is available to prospects under NDA. [P2]

Q4. Customer data is retained for the duration of the contract plus 90 days (contract term + 90 days), after which it is deleted from production systems. [P3]

Q5. The current subprocessor list is published at trust.bonusly.com. Customers receive 30 days notice before a new subprocessor is added. [P4]

Q6. Yes. SAML 2.0 single sign-on is supported (along with SCIM user provisioning), available on the Enterprise plan. [P5]

Q7. Production databases are backed up daily. Backups are encrypted and retained for 35 days. [P6]

Q8. Unanswerable from the provided excerpts — no ISO 27001 status is stated. This would be answered by our certification/attestation documentation (e.g., ISO 27001 certificate or a security certifications summary on the Trust portal).

Q9. Unanswerable from the provided excerpts — no uptime SLA percentage is stated. This would be answered by our Service Level Agreement / MSA terms.

Q10. Unanswerable from the provided excerpts — no HIPAA/BAA commitment is stated. This would be answered by our legal terms or a HIPAA readiness/BAA policy document.

Summary: Q1–Q7 answered from policy excerpts (P1–P6, P5 for Q6). Q8–Q10 are not addressed in the provided excerpts and are marked unanswerable rather than guessed.
ops-maintenance 5 tests
skill-file-reconciliationPASS 1.00
Plain-English question: Audit a skill set against its manifest and flag every conflict with a severity and an action.
Prompt — what the model saw
Reconcile the skill set below against its manifest. Find and report (with severity CRITICAL/WARNING/INFO and an action type TRIM_DESC, UPDATE_BODY, DELETE_SKILL, MERGE, or REVIEW): (1) any two skills whose ALWAYS-trigger phrases overlap/duplicate, (2) any circular delegation chain between skills (name it), (3) any delegation target that does not exist (dangling, name it), (4) any version conflict and which skill should survive, (5) how many manifest descriptions exceed 1,024 characters, (6) any hardcoded page ids, dates, or person names in skill bodies, (7) manifest drift in BOTH directions (files with no manifest row, and manifest rows with no file). Make one proposal per finding; rewrite nothing.
cost $0.3813 · 593s · in 64,620 / out 28,726 tok
✓ duplicate_trigger_found✓ circular_chain✓ dangling_target✓ oversized_desc_count✓ hardcoded_values✓ manifest_drift✓ severity_and_action

✓ Correct Answer (ground truth)

Plain-English answer: A severity-ranked list of TRIM_DESC / UPDATE_BODY / DELETE_SKILL / MERGE / REVIEW actions covering every rule violation found.
The answer in one line: files_audited = 14; rank = CRITICAL / WARNING / INFO with one fix action per issue
files_audited14
findingsduplicate trigger, circular delegation chain, dangling file reference, oversized description, hardcoded values, manifest drift
rankCRITICAL / WARNING / INFO with one fix action per issue
graded checksduplicate_trigger_found · circular_chain · dangling_target · oversized_desc_count · hardcoded_values · manifest_drift · severity_and_action

Model's actual answer

show the model's full answer
RECONCILIATION — 14 SKILL.md files vs skill_manifest.csv (14 rows)

Scope arithmetic: manifest rows = 14; SKILL.md files provided = 14; skill_manifest.csv is the manifest itself, not a skill.

=====================================================
(1) ALWAYS-TRIGGER OVERLAP / DUPLICATION
=====================================================

F1a. comms-drafter vs email-drafter — CRITICAL — MERGE
Five trigger phrases appear verbatim in BOTH descriptions: "write me an email", "draft a follow-up", "what should I say", "bump email", "contract nudge" — plus near-verbatim "help me reply" (comms-drafter) / "help me reply to this" (email-drafter). Both also claim the paste-a-message-review/rewrite case. Every email-drafting request satisfies both ALWAYS conditions; routing is non-deterministic.
Proposal: MERGE email-drafter into comms-drafter. comms-drafter is the functional superset (email + Intercom/support + partner/channel). Before retirement, absorb email-drafter's unique body content (Gmail signature retrieval protocol, no-markdown-in-email-body rule) into comms-drafter and repoint deal-strategy-coach's email-drafter reference (see F2).

F1b. pipeline-intelligence-report vs weekly-pipeline-report — CRITICAL — REVIEW
Three overlapping trigger families: (a) pipeline-intelligence-report triggers on "pipeline update"; weekly-pipeline-report triggers on "run the pipeline update" and "update the pipeline" — an ask like "run the pipeline update" matches both. (b) "run the pipeline report" (PIR) vs "generate the pipeline report" / "do the pipeline report" (WPR). (c) "what's the pipeline look like" (PIR) vs "what does pipeline look like" (WPR). Unlike next-to-close, which explicitly disambiguates against pipeline-intelligence-report ("Do not use pipeline-intelligence-report for this... Delegate to pipeline-intelligence-report if the user wants the full scored pipeline instead"), neither of these two carves exclusive vocabulary.
Proposal: REVIEW — assign exclusive trigger vocabulary: pipeline-intelligence-report owns scored/tiered/full-pipeline asks; weekly-pipeline-report owns weekly/SQM/SQO/DS2/bookings-MTD asks; strike the shared phrases from one description each.

F1c. pipeline-intelligence-report vs sales-forecast — WARNING — TRIM_DESC
pipeline-intelligence-report: "Also trigger when Alaina or any VP asks for pipeline health or forecast context." sales-forecast: ALWAYS triggers on "pipeline forecast", "forecast report", "forecast update". A "pipeline forecast" ask matches both.
Proposal: TRIM_DESC — remove "forecast context" from pipeline-intelligence-report's trigger clause; route forecast asks to sales-forecast.

F1d. deal-strategy-coach vs next-to-close — WARNING — TRIM_DESC
deal-strategy-coach triggers when a manager "asks which deals are likely to close"; next-to-close ALWAYS triggers on "which deals are most likely to close". Near-identical.
Proposal: TRIM_DESC — remove the "asks which deals are likely to close" clause from deal-strategy-coach; shortlist asks route to next-to-close.

=====================================================
(2) CIRCULAR DELEGATION CHAIN
=====================================================

F2. deal-strategy-coach → email-drafter → deal-strategy-coach — WARNING — UPDATE_BODY
Evidence: deal-strategy-coach: "When drafting manager-to-prospect emails, use the `email-drafter` skill which automatically retrieves your Gmail signature and appends it to all prospect-facing emails." email-drafter: "For deal strategy, diagnosis, or coaching (not email drafting), use deal-strategy-coach instead." A hybrid ask ("coach this stalled deal and draft the manager email") ping-pongs between the two. comms-drafter feeds the same cycle ("For deep deal strategy, use deal-strategy-coach — this skill drafts, that skill diagnoses").
Proposal: UPDATE_BODY — break one edge of the cycle: either remove deal-strategy-coach's delegation to email-drafter (inline the signature step) or make email-drafter's deal-strategy-coach pointer a one-way lane marker that never re-invokes. If F1a's MERGE is adopted, this cycle collapses once deal-strategy-coach's reference is repointed.

=====================================================
(3) DANGLING DELEGATION TARGETS
=====================================================

F3. Referenced as skill handoffs but present in neither the file set nor the manifest — WARNING — REVIEW
- bonusly-brand (comms-drafter, email-drafter, sales-forecast, signalforge-claim-compressor)
- prospect-research-multithreading (comms-drafter, email-drafter, deal-strategy-coach)
- skill-orchestrator (signalforge-feedback: "Skill registered in skill-orchestrator as a terminal step"; analysis-validator §11)
- signalforge-reports org skill (pipeline-intelligence-report, weekly-pipeline-report — /mnt/skills/organization/signalforge-reports/SKILL.md)
- analysis-validator §12.4 specialists: bonusly-data-questions, bonusly-product-questions, bonusly-business-reporting-questions, bonusly-rewards-questions, bonusly-ppp-questions, bonusly-feature-flag-questions, bonusly-deal-desk-questions, bonusly-datadog-questions
- analysis-validator §11 cascade files: CUSTOMER_DATA_REFERENCE, HUBSPOT_CONNECTOR_REFERENCE, SIGNALFORGE_PRODUCT_INSIGHT_SKILL
- (file-path dependencies outside the set: /mnt/skills/public/xlsx/scripts/recalc.py in stale-pipeline-report)
Related, unresolvable from provided data: model-selection claims "This skill runs FIRST — before any other skill" — if skill-orchestrator exists in the parent library, its first-position claim cannot be reconciled against it here.
Proposal: REVIEW — verify each target exists in the parent skill library; add manifest rows if in scope, or annotate the references as external.

=====================================================
(4) VERSION CONFLICTS
=====================================================

F4a. analysis-validator internal version conflict — WARNING — UPDATE_BODY
Header/changelog declare v3.6 ("Last Updated: May 9, 2026 (v3.6 — G2-F...)"), and pipeline-intelligence-report's footer cites "Analysis Validator v3.6" — but analysis-validator's own §7 validation-trail template hardcodes "Validator: analysis-validator v3.2".
Survivor: v3.6 (header + changelog + the one external citation all agree; only the §7 template is stale).
Proposal: UPDATE_BODY — correct the §7 trail template to v3.6 or make it version-agnostic.

F4b. pipeline-intelligence-report "v6" vs "v4 Component Vocabulary" — INFO — REVIEW
Frontmatter and title declare "v6 · May 2026"; the component section is titled "v4 Component Vocabulary". If "v4" is the signalforge-reports design-system version, it is intentional but unverifiable here (signalforge-reports is absent — see F3).
Proposal: REVIEW — confirm whether "v4" is the design-system version; if it is a stale skill-version label, relabel.

F4c. Generational conflict: comms-drafter vs email-drafter (cross-ref F1a) — which skill survives: comms-drafter. It is the superset (all external communications, not just email). email-drafter's unique assets (Gmail signature protocol, markdown rules) are absorbed before retirement; deal-strategy-coach's reference is repointed. — CRITICAL — MERGE (same proposal as F1a; recorded here to answer "which skill should survive" directly).

=====================================================
(5) DESCRIPTIONS OVER 1,024 CHARACTERS
=====================================================

F5. Count = 0 of 14 — INFO — REVIEW
Arithmetic (manifest description_chars): 656, 656, 676, 708, 762, 792, 897, 945, 962, 965, 996, 1004, 1006, 1006. Max = 1,006 (pipeline-intelligence-report and signalforge-claim-compressor). 1,006 < 1,024 (18 under the limit). Values > 1,024: none.
Caveat: this uses the manifest's declared counts; exact character counts cannot be verified from the rendered file text — that is a data gap, not a confirmed fact.
Proposal: REVIEW — verify description_chars against the actual YAML at manifest refresh time; no trim needed per the data provided.

=====================================================
(6) HARDCODED PAGE IDs, DATES, PERSON NAMES IN BODIES
=====================================================

F6. — WARNING — UPDATE_BODY
Inventory (cited as given):

Person names:
- analysis-validator §12.3: full roster — Bryce Harmon, Hugo Lindqvist, Dana Mercer, Alex Franklin, Cole Ingram, Gavin Porter, Colleen Perry, Ellie Barton, Ashley Reyer, Megan Franz, Elena Sinclair, Youssef Elkhateeb, Amanda Czenkus, Alaina Loori, Shealagh Coughlin, Ben Castelli, Amani Phipps, John Thomas, Yasmin Wahid (each with a hardcoded HubSpot owner ID); §10/G1-K: "Manish or Amani" (Finance escalation). Note the internal contradiction: §8 says "Do not use hardcoded figures" while §12.3 hardcodes a roster "Updated May 4, 2026".
- pipeline-intelligence-report Phase 1: Bryce Harmon, Dana Mercer, Cole Ingram, Alex Franklin, Gavin Porter + owner IDs ("verified May 2026"); trigger clause names "Alaina".
- sales-forecast: "Manager Forecast (Alaina / VP Sales view)"; changelog "Elena → Alaina".
- partner-digest: "Amani's threads", "Owner: Amani Phipps (RevOps / Partnerships)", "Amani reviews the live Confluence page"; partner contacts Kelli, Jen Lee, Hani, Bryce, Sara; Slack ID U03QLMBL7AR.
- weekly-pipeline-report: "Ben Lavin · Demand Generation" in the H1; "for Ben's review"; "deliver the HTML file to Ben".
- deal-strategy-coach: "routed to Farid for manual qualification".
- signalforge-feedback: example "Gavin Porter Rep Diagnostic".
- signalforge-claim-compressor: attribution "JuliusBrussee/caveman".

Page/document IDs:
- partner-digest: cloudId 73fe98de-a4a3-4869-9f8a-bb1eeed4cf7f, spaceId 1958248479, folder 2286616609, page 2286321666, pages 2265382925, 2236940297, 2237825028, 2239365136, 2238283777.
- sales-forecast: spaceId 2232811524, cloudId 73fe98de-a4a3-4869-9f8a-bb1eeed4cf7f, parent page 2232582148.
- signalforge-feedback: page 2295136266, parent 2234417154, Build Log page 2247295002, spaceId 2232811524.
- deal-strategy-coach: Confluence page 2257879045 (AE Excellence Playbook URL).
- weekly-pipeline-report: spreadsheet IDs 1CLZeOsElVDF_LF0ZG_t2nfwvhnZ6bpwqM_nX3WEYzcw and 1ENuaEcCuLjdKhMvp8FK3Ys1ek5Aw9ZuOZhsHJJFoB_k.
- (Related non-page IDs: HubSpot org ID 1973303 in next-to-close, pipeline-intelligence-report, stale-pipeline-report; Slack channel C0561C1JCPJ in stale-pipeline-report.)

Dates:
- analysis-validator: April 26 2026, May 9 2026, May 4 2026 (x3), March 28 2023, "as of May 2026" expected ranges.
- closed-lost-analysis: "30-deal AI-field sample from May 2026", "Lost to Motivosity on this in May 2026 (Softheon)", "MinIO: rep vacation May 4–12", "confirmed May 2026" (x2).
- model-selection: last_checked 2026-05-19, changelog 2026-05-19, "April 14, 2026" deprecation.
- partner-digest: "May 16, 2026", "May 19, 2026", "June 2, 2026" examples, changelog 2026-05-17.
- sales-forecast: "July 9, 2026" example, changelog April 27, 2026.
- signalforge-claim-compressor: changelog 2026-05-09.
- stale-pipeline-report: changelog 2026-06-10; example dates "5/15", "5/19", "5/7".
- weekly-pipeline-report: "Q2 (April 1 – June 30, 2026)"; Q1 2026 static actuals ($365,152 vs $475,000 plan; $2,490,532 vs $3,288,000 forecast).
- pipeline-intelligence-report: "verified May 2026", "last modified March 2023".

Clean on all three categories: comms-drafter, email-drafter. Contrast: stale-pipeline-report's own design rule is "Never hardcode rep names or owner IDs" — the roster hardcoding in analysis-validator and pipeline-intelligence-report violates the same principle the set elsewhere enforces.
Proposal: UPDATE_BODY — externalize rosters, owner IDs, Confluence page/space IDs, and sample dates to runtime lookups or a shared reference file; retain only declared system constants (e.g., stage IDs, which pipeline-intelligence-report already labels "verify at run time").

=====================================================
(7) MANIFEST DRIFT — BOTH DIRECTIONS
=====================================================

F7. Zero drift in either direction — INFO — REVIEW
Arithmetic: manifest rows = 14; SKILL.md files = 14. Files with no manifest row: 14 − 14 = 0. Manifest rows with no file: 14 − 14 = 0. All 14 declared_name values match their file's frontmatter name exactly (verified pairwise).
Caveat: description_chars and body_lines cannot be verified against the rendered text — count-level drift is a data gap; presence/absence reconciliation is complete and clean.
Proposal: REVIEW — add automated description_chars/body_lines verification to the manifest refresh, since those two columns are unverified by this reconciliation.

=====================================================
ADDITIONAL FINDINGS (outside the 7 requested categories — surfaced during reconciliation)
=====================================================

F8. stale-pipeline-report internal contradiction — WARNING — UPDATE_BODY
Phase 4: "ZoomInfo (ZoomInfo:account_research) — run for every deal, regardless of ARR or stage." Execution Sequence step 4: "Enrich thin deals — ZoomInfo:account_research and/or web_search for DS3+ or ARR ≥ $20K with empty ai_why_buys." Mutually exclusive scopes.
Proposal: UPDATE_BODY — pick one scope and align both passages.

F9. Terminal-position conflict: analysis-validator vs signalforge-claim-compressor — INFO — REVIEW
analysis-validator: "It is the last thing that runs before any output is published, shared, or formatted" (Step N-1). signalforge-claim-compressor: "This skill runs after all data has been gathered and after analysis-validator has run... data pull → analysis → analysis-validator (tiers assigned) → claim-compressor → output" and calls itself "The final style gate before any SignalForge output ships." Both claim the last pre-output position.
Proposal: REVIEW — designate one authoritative execution order and state it identically in both skills (signalforge-feedback already implies validator → compressor → feedback).

Nothing was rewritten; all items above are findings plus one proposal each.
runbook-from-slack-threadPASS 1.00
Plain-English question: Turn the resolved incident thread into a runbook.
Prompt — what the model saw
Turn the resolved incident thread into a runbook: ordered steps with the exact command or action taken, who ran it, how success was verified, and rollback for each step that changed state. Every step must trace to a message in the thread. Add nothing the thread does not contain; where the thread is unclear, mark the step as needing confirmation.
cost $0.0393 · 85s · in 0 / out 4,095 tok
✓ ordered_steps✓ rollback✓ needs_confirmation✓ sourced

✓ Correct Answer (ground truth)

Plain-English answer: Ordered steps from the thread, each with the exact command, who ran it, how success was verified, and rollback where state changed.
The answer in one line: rule = no steps beyond what the thread contains; every step traces to a message id.
ambiguous_steporder=4; msg=M05; mark=needs-confirmation
ruleno steps beyond what the thread contains; every step traces to a message id.
stepsorder=1; msg=M01; actor=Bryce Harmon; action=acknowledge PagerDuty alert, take IC; state_change=False; verified_by=M01 | order=2; msg=M02; actor=Farid Osman; action=check queue depth; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M02 (48,213 pending) | order=3; msg=M04; actor=Farid Osman; action=pause auto-enqueue via feature flag; command=bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'; state_change=True; rollback=bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'; verified_by=M07 (queue depth falling) | order=4; msg=M05; actor=Elena Sinclair; action=cleared the dead set; state_change=True; needs_confirmation=thread gives no command or method for clearing the dead set | order=5; msg=M06; actor=Bryce Harmon; action=scale reward-worker to 6 replicas; command=kubectl scale deployment/reward-worker --replicas=6; state_change=True; rollback=kubectl scale deployment/reward-worker --replicas=3; verified_by=M07/M08 | order=6; msg=M08; actor=Cole Ingram; action=verify queue drained and error rate baseline; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M08 (returns 0) … (+2 more)
graded checksordered_steps · rollback · needs_confirmation · sourced

Model's actual answer

show the model's full answer
RUNBOOK — reward-worker queue backlog, 2026-08-29 (#incident-2026-08-29-reward-queue, resolved)
All steps trace to thread messages [M##]. Derived arithmetic is labeled as derived.

STEP 1 — Acknowledge alert, take IC
  Who: Bryce Harmon [M01, 14:02:10Z]
  Action: Acknowledge PagerDuty alert (reward-worker queue depth > 10k); take IC role.
  Command: not stated in thread — needs confirmation.
  Verification: not stated — needs confirmation.
  Rollback: n/a (no state change).

STEP 2 — Measure queue depth
  Who: Farid Osman [M02, 14:04:33Z]
  Command: bundle exec rake sidekiq:queue_depth
  Result: 48,213 pending jobs. Normal cited as under 500 (per Farid Osman).
  Verification: command output itself (read-only).
  Rollback: n/a.

STEP 3 — Inspect dead set
  Who: Farid Osman [M03, 14:06:02Z]
  Action: inspect Sidekiq dead set; exact command not stated in thread — needs confirmation.
  Result: 112 jobs, all Redis::TimeoutError, timestamped around 13:58.
  Verification: inspection output (read-only).
  Rollback: n/a.

STEP 4 — Pause enqueue (STATE CHANGE)
  Who: Farid Osman [M04, 14:08:45Z]
  Command: bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'
  Verification: not explicitly verified in thread — needs confirmation. (Enqueue remained off until Step 8.)
  Rollback: bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)' — stated in thread; executed in Step 8.

STEP 5 — Clear dead set (STATE CHANGE) — NEEDS CONFIRMATION
  Who: Elena Sinclair [M05, 14:15:20Z]
  Action: "cleared out the dead set" while in the console. Exact command not in thread — needs confirmation.
  Verification: none stated — needs confirmation.
  Rollback: none documented — needs confirmation.
  Note: thread does not state whether the 112 dead jobs from Step 3 were retried, re-enqueued, or discarded.

STEP 6 — Scale workers up (STATE CHANGE)
  Who: Bryce Harmon [M06, 14:21:07Z]
  Command: kubectl scale deployment/reward-worker --replicas=6 (was 3)
  Arithmetic (derived): 3 -> 6 replicas = +3 replicas, 2x capacity.
  Verification: queue depth falling afterward [M07].
  Rollback: kubectl scale deployment/reward-worker --replicas=3 — stated in thread; executed in Step 9.

STEP 7 — Verify drain (read-only)
  Who: Farid Osman [M07, 14:33:41Z], then Cole Ingram [M08, 14:47:55Z]
  Action/results: Farid Osman — queue depth 9,400, falling ~1,200/min. Cole Ingram — bundle exec rake sidekiq:queue_depth returns 0; Datadog error rate back to baseline.
  Rollback: n/a.

STEP 8 — Re-enable enqueue (STATE CHANGE)
  Who: Bryce Harmon [M09, 14:49:10Z]
  Command: bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'
  Verification: 40 new jobs processed cleanly in the next 3 minutes (per Bryce Harmon).
  Rollback: bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)' (same command as Step 4).

STEP 9 — Scale workers back down (STATE CHANGE)
  Who: Bryce Harmon [M10, 14:55:00Z]
  Command: kubectl scale deployment/reward-worker --replicas=3
  Verification: queue stable at 0 (per Bryce Harmon). Incident declared resolved.
  Rollback: kubectl scale deployment/reward-worker --replicas=6 (same command as Step 6).

DERIVED FIGURES (arithmetic from thread numbers only; none of these rates are stated in the thread)
  - Total drain: 48,213 (14:04:33Z) -> 0 (14:47:55Z) = 48,213 jobs in 43 min 22 s.
  - 48,213 - 9,400 = 38,813 jobs between 14:04:33Z and 14:33:41Z (29 min 8 s) = ~1,338 jobs/min average. Caveat: replicas went 3 -> 6 at 14:21:07Z mid-interval, so this average mixes both capacities.
  - 9,400 jobs between 14:33:41Z and 14:47:55Z (14 min 14 s) = ~661 jobs/min average, below the ~1,200/min rate Farid Osman reported at 14:33:41Z. Discrepancy unexplained in thread — needs confirmation if drain rate matters.
  - Enqueue paused 14:08:45Z -> 14:49:10Z = 40 min 25 s.

GAPS EXPLICITLY MISSING FROM THREAD
  1. Step 1: acknowledgment command/verification not captured.
  2. Step 3: dead-set inspection command not captured.
  3. Step 5: exact console action, verification, rollback, and fate of the 112 dead jobs all undocumented — needs confirmation.
  4. Root cause of the Redis::TimeoutError cluster around 13:58 is not identified anywhere in the thread.
datadog-incident-timelinePASS 1.00
Plain-English question: Sequence the failure from the logs — what broke first, and what cascaded after it?
Prompt — what the model saw
From the log slice, sequence the failure: first error with timestamp, the cascade in order, the service and job involved, and the Datadog query that would confirm the first error. State what the logs do not show.
cost $0.0373 · 72s · in 617 / out 3,398 tok
✓ first_error_ts✓ service_named✓ cascade_ordered✓ query✓ gaps_stated

✓ Correct Answer (ground truth)

Plain-English answer: A reward-service Redis timeout at 14:01:12Z cascades to sidekiq at 14:01:40 and api-gateway at 14:03:05, with the Datadog query that confirms the first error.
The answer in one line: job_class_involved = RewardGiveJob
cascade_orderservice=reward-service; first_error_at=2026-09-03T14:01:12Z | service=sidekiq; first_error_at=2026-09-03T14:01:40Z | service=api-gateway; first_error_at=2026-09-03T14:03:05Z | service=web-app; first_error_at=2026-09-03T14:03:30Z
clean_servicespostgres
confirming_queryservice:reward-service status:error "@message:Redis::TimeoutError"
first_errortimestamp=2026-09-03T14:01:12Z; service=reward-service; message=Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s
job_class_involvedRewardGiveJob
logs_do_not_showroot cause of the Redis timeout, redis-server's own logs, infrastructure metrics (CPU/memory/network), any deploy or config change preceding the incident, who applied mitigation or when (no actor fields in logs)
graded checksfirst_error_ts · service_named · cascade_ordered · query · gaps_stated

Model's actual answer

show the model's full answer
FAILURE SEQUENCE

First error
2026-09-03T14:01:12Z — reward-service — ERROR — "Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s"
This is the root failure in the slice: reward-service lost its connection to Redis at redis-primary:6379. Last healthy reward-service log before it: 13:59:30Z "job enqueued" (gap of 14:01:12 − 13:59:30 = 1 min 42 s, unexplained by the logs).

Cascade in order
1. 14:01:12Z — reward-service: Redis connection timeout to redis-primary:6379 (root)
2. 14:01:20Z, 14:01:30Z, 14:01:40Z — reward-service: "retry exhausted for RewardGiveJob" (3 retries, 10 s apart: 20→30→40)
3. 14:01:40Z — sidekiq: "RewardGiveJob failed: Redis::TimeoutError; retrying in 60s"
4. 14:02:28Z — sidekiq: RewardGiveJob failed again, retrying
5. 14:02:30Z — sidekiq WARN: "Queue reward depth above 10,000" (backlog building)
6. 14:03:05Z — api-gateway: "502 upstream timeout calling reward-service /gives"
7. 14:03:30Z — web-app: "Give form submission failed: upstream 502 from api-gateway" — first user-facing impact, 14:03:30 − 14:01:12 = 2 min 18 s after the root error
8. 14:03:31Z–14:06:52Z — repeating loop: sidekiq job failures, api-gateway 502s, web-app submission failures
9. 14:22:10Z — reward-service: "Redis connection restored; resuming job processing" — recovery, 20 min 58 s after first error
10. 14:24:45Z — sidekiq: "Queue reward depth below 500" — backlog drained, 23 min 33 s after first error (22 min 15 s after the >10,000 warning)

Service and job involved
- Originating service: reward-service. Downstream services affected in order: sidekiq → api-gateway → web-app.
- Job: RewardGiveJob (per logs). sidekiq_jobs.csv shows 12 RewardGiveJob failures (J-00001 through J-00012) plus 4 RecognitionDigestJob failures (J-00013 through J-00016) = 16 total.

Datadog query to confirm the first error
service:reward-service status:error "Connection to Redis at redis-primary:6379 timed out"
Time window: 2026-09-03T14:00:00Z to 14:05:00Z. Broader confirmation of the error class: service:reward-service status:error "Redis::TimeoutError".

What the logs do not show
- The cause of the Redis timeout. No Redis-server logs, no failover/OOM/network/deploy events — nothing explains why redis-primary:6379 became unreachable.
- Anything between 13:59:30Z and 14:01:12Z that would bridge healthy state to failure.
- RecognitionDigestJob failures. The 4 RecognitionDigestJob failures in sidekiq_jobs.csv (J-00013 to J-00016, 14:02:36Z–14:05:50Z) never appear in the log slice — the logs name only RewardGiveJob.
- Job-level detail in the logs: no job IDs. The CSV's earliest failure is J-00005 at 14:01:46Z, 6 s after sidekiq's 14:01:40Z failure log — the logs don't show which job that was, and the CSV doesn't cover it.
- Whether failed RewardGiveJob/RecognitionDigestJob instances were eventually delivered after 14:22:10Z — no job-success logs follow the recovery line.
- Whether web-app submissions that received 502s were lost or user-retried.
- Actual queue depth values — only the thresholds ">10,000" (14:02:30Z) and "<500" (14:24:45Z) appear; no measurements in between.
feature-flag-state-summaryPASS 1.00
Plain-English question: Summarize what each feature flag controls and who has it on.
Prompt — what the model saw
Summarize the state of the listed feature flags: what each controls per the code excerpt, which companies or segments have it on, targeting rules, and any flag with no code reference or no targets. Company counts must match the export. Do not describe a flag that is not in the export.
cost $0.0247 · 36s · in 0 / out 1,681 tok
✓ orphans✓ no_targets✓ counts_match✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: Recognition streaks on for 42 beta companies, redeem-flow redesign off with a 12-company target list, plus any flag with no code reference or no targets.
The answer in one line: rule = company counts must equal the export; do not describe flags absent from the export.
code_referencedrecognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, redeem_flow_redesign, analytics_dashboard_v3, ms_teams_app_v2
flags{"recognition_streaks_v2": {"state": "on", "targeting_rules": "segment:beta_companies", "company_count": 42}, "points_budget_guardrails": {"state": "on", "targeting_rules": "all_companies", "company_c
flags_with_no_targetspaused_offboard_cleanup
orphan_flags_no_code_referencelegacy_give_modal, survey_boosters_q3, paused_offboard_cleanup
rulecompany counts must equal the export; do not describe flags absent from the export.
graded checksorphans · no_targets · counts_match · no_invented

Model's actual answer

show the model's full answer
FEATURE FLAG SUMMARY (source: flags_export.csv + flag_code.md only)

Flags with code references:

1. recognition_streaks_v2 — ON
   Controls: In app/models/recognition.rb, gates StreakTracker.record(give) — i.e., streak tracking on recognition gives.
   Targeting: segment:beta_companies — 42 companies.

2. points_budget_guardrails — ON
   Controls: In app/services/budget_service.rb, gates BudgetService.new(company).enforce!(giver, points) — budget/guardrail enforcement on point gives.
   Targeting: all_companies — 220 companies. (The export does not state the total company base separately, so I cannot confirm 220 is the entire customer base, only that this flag's count is 220.)

3. slack_dm_nudges — ON
   Controls: In app/jobs/nudge_job.rb, the job exits early unless enabled; when enabled, sends SlackDm.send_nudge(user).
   Targeting: segment:region_na — 87 companies.

4. redeem_flow_redesign — OFF
   Controls: In app/controllers/redeem_controller.rb, switches redemption UI: enabled renders RedeemV2Component, otherwise RedeemV1Component.
   Targeting: targeted_list — 12 companies. State is off, so per this export the flag is not enabled even for the listed 12.

5. analytics_dashboard_v3 — ON
   Controls: In app/controllers/analytics_controller.rb, gates @dashboard = AnalyticsV3.new(company) — the v3 analytics dashboard.
   Targeting: segment:tier_three — 65 companies.

6. ms_teams_app_v2 — OFF
   Controls: In app/services/teams_installer.rb, gates TeamsAppV2.install(company).
   Targeting: targeted_list — 9 companies. State is off, so the redesign is not active per this export.

Flags with NO code reference in the excerpt:

7. legacy_give_modal — OFF. Targeting: segment:legacy_plan — 14 companies. What it controls is not determinable from the provided excerpt; no matching FeatureFlags.enabled? call is shown.

8. survey_boosters_q3 — ON. Targeting: segment:legacy_plan — 7 companies. No code reference in the excerpt; behavior unknown from this data. Note: this shares the segment:legacy_plan rule with legacy_give_modal, but the export gives per-flag counts (14 vs 7) — I cannot confirm whether these are overlapping or distinct companies.

9. paused_offboard_cleanup — OFF. No targeting rules and 0 companies — this is the only flag with no targets. Also no code reference in the excerpt.

Count check (sum of per-flag company_count):
42 + 220 + 87 + 12 + 65 + 9 + 14 + 7 + 0 = 456
This is a sum of per-flag counts, not unique companies — the export provides no way to deduplicate companies appearing in multiple segments/lists.

Caveats, stated explicitly:
- Three of nine flags (legacy_give_modal, survey_boosters_q3, paused_offboard_cleanup) have no code reference in the excerpt, so their functional impact cannot be described.
- The export does not identify the companies behind targeted_list (redeem_flow_redesign: 12; ms_teams_app_v2: 9) — no company aliases are provided, so none can be cited.
- Two flags are state=off yet carry a targeted_list (12 and 9 companies); the export does not explain whether those lists are staged rollout candidates or stale.
nda-triagePASS 1.00
Plain-English question: Triage each NDA green, yellow, or red — and quote the clause that decides it.
Prompt — what the model saw
Triage each NDA as GREEN (standard approval), YELLOW (counsel review), or RED (full legal review). Cite the clause that drives each classification by quoting it. Catch embedded non-solicits, non-competes, missing carve-outs, and unusual term or governing law. One line of reasoning per NDA.
cost $0.0193 · 18s · in 0 / out 769 tok
✓ green✓ yellow✓ red✓ non_solicit_caught✓ clause_quoted✓ correct_labels

✓ Correct Answer (ground truth)

Plain-English answer: NDA-1 green (standard mutual terms), NDA-2 yellow for an embedded 18-month non-solicit, with the driving clause quoted for every classification.
The answer in one line: non_solicit_in = NDA-2
checksclassification_per_nda, clause_quotes_exist_in_document, non_solicit_caught, carveouts_checked
classifications{"NDA-1": {"classification": "GREEN", "driving_reason": "standard mutual NDA: 2-year term, standard carve-outs, Delaware law", "driving_clause_quote": "Carve-outs: information that (a) is or becomes p
non_solicit_inNDA-2
graded checksgreen · yellow · red · non_solicit_caught · clause_quoted · correct_labels

Model's actual answer

show the model's full answer
NDA-1 — GREEN
Reasoning: All elements are standard — mutual flow, 2-year term + 3-year survival (5 years total obligation), full carve-out set (a)–(e), Delaware law, express no-license/no-exclusivity; nothing to escalate.

NDA-2 — YELLOW
Reasoning: Embedded non-solicit buried in an otherwise standard NDA — "During the term of this Agreement and for eighteen (18) months thereafter, neither party shall solicit for employment or hire any employee of the other party with whom it came into contact under this Agreement" — the tail extends 18 months past the 3-year term (4.5 years of hiring exposure), so counsel should review scope before signature.

NDA-3 — RED
Reasoning: Four compounding problems — (1) one-way confidentiality (Bonusly discloses, receives no protection); (2) an embedded non-compete barring Bonusly from its own market: "Recipient shall not, directly or indirectly, engage in or provide services to any business that competes with Discloser's business" for 3 years; (3) no confidentiality carve-outs stated at all, so even publicly available or independently developed information could stay technically "confidential" for the full 5-year term; and (4) unusual governing law with exclusive foreign jurisdiction: "the laws of the Republic of Ireland, and the parties submit to the exclusive jurisdiction of its courts" — full legal review required.

Arithmetic summary:
- NDA-1: 2-yr term + 3-yr survival = 5 yrs max obligation
- NDA-2: 3-yr term + 18-mo non-solicit tail = 4.5 yrs hiring restriction
- NDA-3: 5-yr term with no carve-outs + 3-yr non-compete running from Effective Date

Note: Party identities are redacted in the source file, so I cannot assess whether Party C's "business" overlaps Bonusly's — that's a missing data point legal should resolve, since the breadth of the non-compete depends on it.