← back to all models ← RevenueBench home
deepseek

deepseek/deepseek-flash-latest

Mean 0.956 · 32/40 perfect tests · $0.96 total · modalities in: text, image · out: text · each card: the prompt → the correct answer (gold) → this model's actual answer

What do the modalities mean?

Modalities in means what you can send this model: text, images, files, video, or audio. Modalities out means what it can send back. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what this model can handle in real work.

data-crm 5 tests
stage-hygiene-auditPASS 1.00
Plain-English question: How much of the open DS2+ pipeline is missing a substantive why-buys, and is anything missing DS1 timestamps or out of order?
Prompt — what the model saw
Audit stage hygiene on the open pipeline below. Rules: a deal at DS2 or later must have a substantive why-buys (15+ characters, not a placeholder such as tbd, n/a, see notes). A why_buys value of the form LONG(n) means the real text was present but truncated for display; treat it as present and substantive. Every deal should have a DS1 entry timestamp. A deal whose stage entry timestamps are out of order regressed (an earlier-stage timestamp is later than a later-stage timestamp). Report: percent of DS2+ deals non-compliant on why-buys overall and by owner (use the owners table for names), count of deals with no DS1 timestamp, count that regressed, and three process fixes. State definitions before numbers.
cost $0.0160 · 270s · in 20,154 / out 25,035 tok
✓ noncomp_pct✓ no_ds1_0✓ regressed_0✓ per_owner_breakdown✓ definitions_first✓ three_fixes

✓ Correct Answer (ground truth)

Plain-English answer: 81.5% of the 124 DS2+ deals are missing a substantive why-buys — Alex Franklin 48, Cole Ingram 19, Dana Mercer 18, Bryce Harmon 15, Farid Osman 1 — and every deal has a DS1 timestamp with none regressed.
The answer in one line: n_deals = 156; ds2plus = 124; noncomp_pct = 81.5; bryce_noncomp = 15; no_ds1 = 0; regressed = 0 (+1 more below)
n_deals156
bryce_noncomp15
ds2plus124
no_ds10
noncomp_pct81.5
regressed0
per_owner_noncompAlex Franklin=48 of 61 (78.7%); Cole Ingram=19 of 22 (86.4%); Dana Mercer=18 of 22 (81.8%); Bryce Harmon=15 of 18 (83.3%); Farid Osman=1 of 1 (100%)
graded checksnoncomp_pct · no_ds1_0 · regressed_0 · per_owner_breakdown · definitions_first · three_fixes

Model's actual answer

show the model's full answer
DEFINITIONS
- Scope: "DS2+" = any deal whose current stage is DS2, DS3, DS4, or DS5 (124 of 156 open deals).
- why-buys substantive: non-empty text of 15+ characters; LONG(n) counts as present/substantive; bare placeholders (tbd, n/a, see notes, etc.) do not.
- DS1 timestamp present: t_ds1 field is non-empty.
- Regressed: any earlier-stage entry timestamp is strictly later than a later-stage entry timestamp among the populated t_ds1…t_ds5 fields (e.g. t_ds2 > t_ds3 is fine; t_ds2 < t_ds3 is a regression).

RESULTS

1. DS2+ why-buys non-compliance — overall
   Non-compliant / DS2+ = 101 / 124 = 81.5%

   By owner:
   - Alex Franklin: 48 / 61 = 78.7%
   - Bryce Harmon: 15 / 18 = 83.3%
   - Dana Mercer: 18 / 22 = 81.8%
   - Cole Ingram: 19 / 22 = 86.4%
   - Farid Osman: 1 / 1 = 100.0% (Deal-499BF6, the only DS2+ deal he owns)
   - Elena Sinclair: 0 DS2+ deals (her only open deal, Deal-57FF13, is DS1) — not rateable
   Check: 18+61+22+22+1 = 124 ✓

2. Deals with no DS1 entry timestamp: 0
   Every one of the 156 open deals has a populated t_ds1.

3. Deals with out-of-order (regressed) stage timestamps: 0
   All populated t_ds1…t_ds5 chains are monotonically non-decreasing.
   Data note (not a violation): the audit found no evidence of backdating, though 11 deals have a current-stage timestamp absent or earlier than adjacent fields in odd ways (e.g., Deal-317E6F at DS3 has no t_ds4; several DS5 deals like Deal-D348E1 carry no t_ds5) — stage-entry capture at the top of the funnel is incomplete even though the ordering rule is satisfied.

4. Three process fixes
   a. Gate DS2→DS3 advancement on a why-buys field that is required and validated (≥15 chars, blocklist of tbd/n/a/see notes). At 81.5% non-compliance this is a field-level enforcement problem, not a coaching problem — the CRM currently lets deals promote with the field blank.
   b. Auto-capture stage-entry timestamps via workflow on stage change (write t_dsN server-side, make prior-stage timestamps immutable) so ordering is machine-enforced; audit the 11 DS5/DS3 deals missing their current-stage timestamp to fix the pipeline-history plumbing.
   c. Weekly hygiene report to owners with a forecast-category link: deals at DS2+ without substantive why-buys cannot be set to COMMIT/BEST_CASE. Commit deals without why-buys (e.g., Deal-C6FE92 at $72,000, Deal-2D1F1B lineage) are the biggest forecasting blind spot — tie field completion to forecast eligibility so the pipeline number is defensible.
owner-resolution0.83
Plain-English question: Every open deal has a blank owner name — resolve each one to a real person using the owners table.
Prompt — what the model saw
The deal extract below has an owner name field that is blank on nearly every row and an owner id field that is populated. Using the owners table, resolve each open deal to an owner name. Report: number of deals resolved, deals whose owner id has no match in the owners table (list them), the owner ids that map to archived/deactivated owners (list their names), and the total pipeline amount per resolved owner. Do not guess a name for any unmatched id.
cost $0.0303 · 456s · in 40,066 / out 41,884 tok
✓ resolved_count✓ unmatched_none✓ deactivated_flagged✓ jake_amount✓ justin_amount✗ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: All 156 deals resolve to six named owners; jake_amount maps to Bryce Harmon ($1,054,144) and justin_amount to Alex Franklin ($624,310).
The answer in one line: jake_amount = 1,054,144; justin_amount = 624,310; resolved = 156
jake_amount1,054,144
justin_amount624,310
resolved156
deactivated
per_ownerBryce Harmon=1,054,144; Alex Franklin=624,310; Dana Mercer=341,195; Cole Ingram=288,161.43; Farid Osman=4,134; Elena Sinclair=2,100
graded checksresolved_count · unmatched_none · deactivated_flagged · jake_amount · justin_amount · no_invented

Model's actual answer

show the model's full answer
All joins computed programmatically (156 deal rows parsed, exact string match on owner_id). Results:

**1. Deals resolved: 156 of 156 (100%)**
Distinct owner ids appearing in deals_open.csv: 6 — 119337721, 84342457, 83155923, 83155924, 716654662, 701163055. All 6 exist in owners.csv.

**2. Unmatched owner ids: NONE**
No deal row carries an owner_id missing from the owners table. (No rows had a blank owner_id either, so nothing falls through.)

**3. Owner ids mapping to archived/deactivated owners: NONE referenced by deals**
The owners table has 2 archived entries — Gavin Porter (1520255671) and Hugo Lindqvist (77260721) — but neither id appears on any deal row, so no open deal resolves to an archived owner.

**4. Total pipeline per resolved owner**

| Owner | owner_id | # deals | Total amount |
|---|---|---|---|
| Bryce Harmon | 119337721 | 35 | $1,054,144.00 |
| Alex Franklin | 84342457 | 67 | $624,310.00 |
| Dana Mercer | 83155923 | 24 | $341,195.00 |
| Cole Ingram | 83155924 | 22 | $288,161.43 |
| Farid Osman | 716654662 | 7 | $4,134.00 |
| Elena Sinclair | 701163055 | 1 | $2,100.00 |

Arithmetic check:
- Deal count: 35 + 67 + 24 + 22 + 7 + 1 = 156 ✓ (equals total rows)
- Amount: 1,054,144.00 + 624,310.00 = 1,678,454.00; + 341,195.00 = 2,019,649.00; + 288,161.43 = 2,307,810.43; + 4,134.00 = 2,311,944.43; + 2,100.00 = **$2,314,044.43** — matches the independent grand total of all 156 amount fields.
- Spot-check Farid Osman: Deal-499BF6 (1,249) + Deal-8BA24E (2,880) + five $1 deals Deal-03BA69, Deal-117863, Deal-F17780, Deal-2B39B0, Deal-8FDCD2 (5) = 1,249 + 2,880 + 5 = 4,134 ✓
- Spot-check Elena Sinclair: single deal Deal-57FF13 = 2,100 ✓

Two data caveats visible in the extract (flagging, not correcting): eight deals carry amount = 1 (likely placeholder), including Deal-012CB1, Deal-483B2D, Deal-3795AD and Farid Osman's five $1 DS1 rows — they are included in the totals above at face value. Also Deal-57FF13 (Elena Sinclair) has a blank last_contacted_field and blank forecast detail on t_ds2+.
gong-hubspot-join-integrityPASS 1.00
Plain-English question: How many call records point at deal ids that do not exist in the open pipeline?
Prompt — what the model saw
Check the integrity of the call-to-deal mapping below. The gong table lists calls in the last 90 days keyed to HubSpot deal ids; deals_open lists the current open default-pipeline deals. Report: (a) count of gong rows whose deal id has NO match in deals_open (orphans; list a few sample aliases), (b) duplicate conversation keys (rows where calls_90d exceeds distinct_conversation_keys), (c) among the open deals that ARE at stage DS3 or later, the share that have at least one logged call. Give sample aliases for the orphan finding, all taken from the data. Note: gong_owner_id is a Gong user id, not a HubSpot owner id.
cost $0.1989 · 72s · in 89,898 / out 19,119 tok
✓ orphan_count✓ dup_keys_0✓ ds3_coverage✓ sample_ids_present

✓ Correct Answer (ground truth)

Plain-English answer: 30 Gong call rows are orphans keyed to dead deal ids, none are duplicates, and DS3 call coverage is 29.4%.
The answer in one line: orphans = 30; ds3_cov_pct = 29.4; dup = 0
orphans30
ds3_cov_pct29.4
dup0
orphan_aliasesDeal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42, Deal-36EA09
graded checksorphan_count · dup_keys_0 · ds3_coverage · sample_ids_present

Model's actual answer

show the model's full answer
Call-to-deal mapping integrity, from data provided only:

(a) Orphans - gong deal id with NO match in deals_open:
67 gong rows - 37 matched = 30 orphans.
30 / 67 = 44.8% orphan.
Sample orphan aliases, all from gong_calls_by_deal_90d.csv: Deal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42.

(b) Duplicate conversation keys - rows where calls_90d > distinct_conversation_keys:
0 / 67 rows.
All 67 rows have calls_90d = distinct_conversation_keys. No duplicates.

(c) Open deals at DS3 or later with at least one logged call:
deals_open rows = 156. DS3 or later (DS3/DS4/DS5) = 85.
Of those 85, deal_id found in gong table = 25.
25 / 85 = 29.4% have at least one logged call. 60 / 85 = 70.6% have zero.

Note: gong_owner_id not joined, per instruction it is a Gong user id, not a HubSpot owner id.
snowflake-sql-trapped-catalogPASS 1.00
Plain-English question: Write one Snowflake query for first-month usage per customer using only the provided, trap-laden catalog.
Prompt — what the model saw
Using ONLY the catalog provided, write one Snowflake SQL query returning, per customer company, for its first calendar month as a customer: unique givers, recognition count, successful redemption count. Read the catalog notes carefully: some tables are stale or unpopulated and must be avoided, and one documented filter (excluding deleted givers) must NOT be applied to historical giving counts. Put a comment above each FROM or JOIN explaining the table choice. Return only fenced SQL.
cost $0.0040 · 134s · in 189 / out 7,781 tok
✓ avoids_stale_deals✓ avoids_unpopulated_live✓ uses_cohort_summary✓ succeeded_filter✓ no_giver_deletion_filter✓ returns_sql

✓ Correct Answer (ground truth)

Plain-English answer: A single correct SQL query that uses the right catalog tables and avoids the trap columns — graded on the exact tables and columns.
The answer in one line: use_table = PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
avoid_stalePRODUCTION.HUBSPOT.DEALS (stale)
avoid_unpopulatedV2_LIVE.OBJECTS_DEALS (unpopulated)
forbidden_filterdeleted-giver exclusion must NOT be applied
outputone fenced SELECT with per-FROM/JOIN table-choice comment
required_filtersucceeded redemptions only
use_tablePRODUCTION.PLG.COMPANY_COHORT_SUMMARY
graded checksavoids_stale_deals · avoids_unpopulated_live · uses_cohort_summary · succeeded_filter · no_giver_deletion_filter · returns_sql

Model's actual answer

show the model's full answer
```sql
-- Per customer company, its FIRST calendar month as a customer:
-- unique givers, recognition count, successful redemption count.
--
-- Data-availability check against the catalog (missing data stated explicitly):
--   * This catalog contains NO recognition/giving event table, NO giver/user
--     table, and NO company-events table with dated gives or givers. Therefore
--     COUNT(DISTINCT giver) and COUNT(recognition) cannot be computed at event
--     level, and the documented "exclude deleted givers" filter — which must
--     NOT be applied to historical giving counts — is not even expressible
--     here. No deleted/archived predicate is added anywhere below.
--   * The only table in the catalog carrying per-company first-month giving
--     metrics is PRODUCTION.PLG.COMPANY_COHORT_SUMMARY (M1_USERS, M1_GIVES,
--     M1_REDEMPTIONS). This query reads those pre-aggregated month-1 values;
--     no row-level arithmetic is possible or needed (metrics are taken
--     verbatim, not summed or re-derived).
--   * MISSING (not invented): the catalog does not name the per-company key
--     column, the M1 anchor date (which of SIGNUP_DATE / ACTIVATED_DATE /
--     FIRST_SUB_PAYMENT_DATE defines month 1), whether M1_USERS counts givers
--     specifically or all active users, and the scope caveat that the table
--     covers SELF-SERVE companies only. The placeholder c.__COMPANY_KEY__ must
--     be replaced with the real key column before running.

SELECT
    c.__COMPANY_KEY__        AS company_key,          -- placeholder: key column undocumented in catalog
    c.CURRENT_STATUS,
    c.SIGNUP_DATE,
    c.ACTIVATED_DATE,
    c.FIRST_SUB_PAYMENT_DATE,
    c.M1_USERS               AS unique_givers_month1,      -- month-1 unique givers, per catalog; NOT cross-validated (no giver-events table exists in this catalog)
    c.M1_GIVES               AS recognition_count_month1,  -- month-1 recognition count
    c.M1_REDEMPTIONS         AS successful_redemptions_month1
-- COMPANY_COHORT_SUMMARY: chosen because it is the only catalog table with
-- per-company first-month gives/givers/redemptions, refreshed as part of the
-- PLG cohort pipeline. Avoided instead:
--   PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 — redemption
--     events, but the catalog documents only its STATE column (no company,
--     giver, or event-date keys; and the literal success STATE value is cut
--     off mid-note as "Only STATE ="), so counting from it would require
--     inventing columns/values. M1_REDEMPTIONS already supplies month-1
--     successful redemptions.
--   HUBSPOT_HUB_1973303.V2_LIVE.OBJECTS_DEALS — unpopulated (do not use).
--   PRODUCTION.HUBSPOT.DEALS — stale, last sync 2023-03 (do not use).
--   PRODUCTION.HUBSPOT.GONG_HUBSPOT_MAP_FAST — retired pattern (do not use).
--   OBJECTS_DEALS / OWNERS / OBJECTS_CONTACTS / HS_ENGAGEMENTS_ENRICHED /
--     GONG_* / ASSOCIATIONS_DEALS_TO_CONTACTS — correct per their notes, but
--     they hold deal, owner, contact, engagement, and call data only; none
--     contain Bonusly gives, givers, or redemptions. OWNERS.ARCHIVED is a
--     HubSpot owner field, not a giver-deleted flag, and is deliberately not
--     joined or filtered on.
FROM PRODUCTION.PLG.COMPANY_COHORT_SUMMARY AS c
-- No WHERE clause: applying any deleted/excluded-population filter here would
-- understate historical giving, which the rules forbid for these counts.
ORDER BY c.__COMPANY_KEY__;
```
crm-completeness-enrichmentPASS 1.00
Plain-English question: Audit the CRM for missing and invalid fields, and propose the fixes.
Prompt — what the model saw
Audit the CRM extract for completeness and propose fixes. Deals need owner, stage, amount, close date, why-buys; companies need industry, employee count, HQ country; contacts need email, title, persona. Report completeness percent per field. Identify duplicate company clusters (name variants, shared domain) with a survivor per cluster. List invalid emails and domain mismatches. Fill missing company fields only where the enrichment export has a matching row; where CRM and enrichment disagree list both and recommend a source. Never invent a value. End with the 10 fixes with the most pipeline amount at stake.
cost $0.1049 · 207s · in 56,230 / out 18,626 tok
✓ dup_clusters✓ invalid_emails✓ survivor_language✓ completeness_pct✓ no_invent

✓ Correct Answer (ground truth)

Plain-English answer: 9 companies missing employee count, 6 missing HQ country, 4 invalid emails, 1 domain mismatch, 2 duplicate clusters, and 25 enrichment matches.
The answer in one line: n_contacts = 52; n_companies = 34; enrichment_matches = 25; invalid_email_count = 4; n_dup_clusters = 2; domain_mismatch_count = 1 (+4 more below)
n_contacts52
domain_mismatch_count1
enrichment_matches25
invalid_email_count4
n_companies34
n_dup_clusters2
company_missingindustry=0; employee_count=9; hq_country=6
domain_mismatchesCT-0011
duplicate_clusters['C-0A092931', 'C-0A092932'], ['C-0A092933', 'C-0A092934']
invalid_emailsCT-0010, CT-0080, CT-0081, CT-0192
graded checksdup_clusters · invalid_emails · survivor_language · completeness_pct · no_invent

Model's actual answer

show the model's full answer
CRM AUDIT - completeness, duplicates, emails, enrichment

DATA SCOPE NOTE
- Files received: companies.csv (34 rows), contacts.csv (52 rows), zoominfo_enrichment.csv (25 rows).
- Deals file NOT provided. Cannot audit deals fields owner / stage / amount / close date / why-buys. Pipeline amount at stake is therefore unknown for all fixes - see last section.

1. COMPLETENESS PERCENT PER FIELD

Companies: denominator 34 for all.
- domain: 34/34 = 100.0%
- industry: 34/34 = 100.0% non-empty, but 9 rows use non-canonical values (tech, Tech with trailing space, health care, SaaS)
- employee_count: 25/34 = 73.5%. Missing 9: C-EC3025, C-96039F, C-44EA29, C-D04904, C-B23205, C-60C75F, C-7BBDFA, C-50D386, C-93C8BF. Arithmetic: 34-9=25; 25/34*100=73.5%
- hq_country: 28/34 = 82.4%. Missing 6: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB. Arithmetic: 34-6=28; 28/34*100=82.4%

Contacts: denominator 52 for all. Covers only 16 of 34 companies. Zero contacts for 14 aliases: C-2C60E5, C-B97B4E, C-2D7423, C-7BBDFA, C-50D386, C-BA969B, C-332637, C-93C8BF, C-EE9FFB, C-C9BB20, C-0A092931, C-0A092932, C-0A092933, C-0A092934.
- email field non-empty: 52/52 = 100.0%; syntactically valid: 48/52 = 92.3%. 4 invalid, see Sec 3.
- title: 39/52 = 75.0%. Missing 13: CT-0000, CT-0022, CT-0072, CT-0080, CT-0081, CT-0092, CT-0120, CT-0121, CT-0122, CT-0132, CT-0141, CT-0162, CT-0170. Arithmetic: 52-13=39; 39/52*100=75.0%
- persona: 37/52 = 71.2%. Missing 15: CT-0000, CT-0022, CT-0041, CT-0060, CT-0070, CT-0081, CT-0082, CT-0092, CT-0110, CT-0132, CT-0162, CT-0171, CT-0172, CT-0180, CT-0181. Arithmetic: 52-15=37; 37/52*100=71.2%

2. DUPLICATE COMPANY CLUSTERS (shared domain)

Cluster A - domain acme-corp.com: C-0A092931 (Technology, 500, US) + C-0A092932 (tech, 510, USA). Survivor recommendation: C-0A092931 - canonical industry and country spelling; employee_count 500 vs 510 conflicts with no enrichment row to arbitrate, requires human verify before merge.
Cluster B - domain globex.io: C-0A092933 (SaaS, 200, US) + C-0A092934 (Technology, 200, US). Survivor recommendation: C-0A092934 - Technology matches house taxonomy used on 32 of 34 rows; SaaS appears once. No enrichment row for globex.io, requires human verify.
No other shared domains. Alias codes give no name-variant signal; no name field to cluster on.

3. INVALID EMAILS AND DOMAIN MISMATCHES

Invalid (fail user@domain.tld pattern), 4 rows:
- CT-0010 (C-66D1FC): user0@
- CT-0080 (C-92D97D): user0@
- CT-0081 (C-92D97D): user1@
- CT-0192 (C-425E2A): user2@
All 4 are truncated at @ with no domain. Fix: re-source address; do not guess.

Domain mismatch (email domain != company domain), 1 row:
- CT-0011 (C-66D1FC): email user1@other-domain.com vs company domain 66d1fc.com and contact domain field 66d1fc.com.
No other mismatches. Contact domain field matches company domain on all 52 rows.

4. STANDARDIZATION ISSUES (do not affect non-empty %, do affect usability)

Industry raw values in CRM: Finance, Healthcare, Manufacturing, Retail, SaaS, Tech with trailing space, Technology, health care, tech.
hq_country raw values in CRM: blank, Canada, UK, US, USA, United States.
Recommend single canonical lists before any merge or reporting.

5. ENRICHMENT FILLS - missing CRM value filled only from matching enrichment row

Enrichment covers 25 domains (2d1f1b.com through 50d386.com). No enrichment row for 9 CRM rows: C-BA969B, C-332637, C-93C8BF, C-EE9FFB, C-C9BB20, C-0A092931, C-0A092932, C-0A092933, C-0A092934. Those cannot be filled - stated explicitly, no value invented.
Fills applied (8 employee_count fills, 0 hq_country fills, 0 industry fills since industry never missing):
- C-EC3025 employee_count = 400 (source zoominfo_enrichment ec3025.com)
- C-96039F employee_count = 400 (source 96039f.com)
- C-44EA29 employee_count = 400 (source 44ea29.com)
- C-D04904 employee_count = 400 (source d04904.com)
- C-B23205 employee_count = 400 (source b23205.com)
- C-60C75F employee_count = 400 (source 60c75f.com)
- C-7BBDFA employee_count = 400 (source 7bbdfa.com)
- C-50D386 employee_count = 400 (source 50d386.com)
hq_country: all 5 missing-HQ rows with enrichment coverage have blank ZI country too (C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5) - no fill possible. C-EE9FFB missing HQ and uncovered - no fill possible.
Post-fill coverage: employee_count 33/34 = 97.1% (only C-93C8BF still missing, uncovered); hq_country stays 28/34 = 82.4%; industry stays 34/34 = 100.0%.

6. CRM VS ENRICHMENT DISAGREEMENTS - both values listed, source recommended

Formatting-only (agree after lower/trim; recommend enrichment spelling as canonical):
- C-66D1FC hq US vs United States; C-950043 US vs United States; C-E51FB7 USA vs United States; C-D0662E US vs United States; C-425E2A USA vs United States; C-2D7423 USA vs United States; C-B23205 US vs United States. Recommend source: enrichment (full country name).
- C-66D1FC industry tech vs Computer Software is substantive, see below. C-425E2A Tech with space vs Computer Software substantive.

Substantive industry (CRM Technology-family vs ZI Computer Software), counts as disagree even though taxonomy differs:
- C-66D1FC: CRM tech vs ZI Computer Software
- C-EC3025: CRM Technology vs ZI Computer Software
- C-44EA29: CRM tech vs ZI Computer Software
- C-92D97D: CRM Technology vs ZI Computer Software
- C-D04904: CRM Technology vs ZI Computer Software
- C-77A95A: CRM Technology vs ZI Computer Software
- C-AA8DDA: CRM Technology vs ZI Computer Software
- C-B25F40: CRM Technology vs ZI Computer Software
- C-60C75F: CRM tech vs ZI Computer Software
- C-425E2A: CRM Tech vs ZI Computer Software
Recommend source: enrichment (ZoomInfo) for industry on these 10, after adopting one taxonomy house-wide; CRM free-text variants tech / Tech are uncontrolled. Exception: where ZI blank, keep CRM (none in this set - all 10 have ZI values).
No substantive employee_count disagreements on covered rows where both present (all match: e.g. C-66D1FC 900=900, C-63A874 340=340). No substantive hq_country disagreements beyond formatting.
Uncovered duplicates cannot be arbitrated: acme 500 vs 510, globex SaaS vs Technology - recommend human verify, no source recommended.

7. TOP 10 FIXES - PIPELINE AMOUNT AT STAKE UNKNOWN

No deals file provided, so pipeline amount per fix cannot be computed. Ranking below is by records affected, as a substitute, highest count first. Amount at stake: data missing for all 10.
1. Standardize industry taxonomy (10 substantive CRM-vs-ZI conflicts + 4 trailing-space/case variants C-425E2A, C-BA969B, C-C9BB20, C-93C8BF + 2 health care C-7BBDFA, C-50D386 + 1 SaaS C-0A092933). Affects 17 rows.
2. Standardize hq_country (US / USA / United States variants across ~12 rows). Affects reporting joins.
3. Fill 8 employee_counts from enrichment (list in Sec 5: C-EC3025, C-96039F, C-44EA29, C-D04904, C-B23205, C-60C75F, C-7BBDFA, C-50D386).
4. Backfill 15 missing contact personas (list in Sec 1).
5. Backfill 13 missing contact titles (list in Sec 1).
6. Repair 4 truncated emails: CT-0010, CT-0080, CT-0081, CT-0192.
7. Resolve CT-0011 domain mismatch (other-domain.com vs 66d1fc.com) - confirm correct person or move contact.
8. Merge Cluster A acme-corp.com (C-0A092931 survivor; resolve 500 vs 510).
9. Merge Cluster B globex.io (C-0A092934 survivor; resolve SaaS vs Technology).
10. Create contacts for 14 contact-less companies, and source hq_country for 6 missing (C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB) plus employee_count for C-93C8BF - none coverable by enrichment provided.
deal-intelligence 5 tests
closed-lost-classificationPASS 1.00
Plain-English question: Why did we lose the last six months of deals — one primary category per deal, and whose side is the reason on?
Prompt — what the model saw
Classify every deal lost in the last 6 months below into one primary category: pricing, competitor, no decision, timing, product gap, champion left, or other. Also assign a side (Bonusly, buyer, unknown). Then summarize: category counts, the side split, the number of deals where the structured closed_lost_tag clearly disagrees with the free-text reason, and the two patterns most worth acting on. Base everything only on the tag and free-text provided.
cost $0.0636 · 106s · in 23,904 / out 7,938 tok
✓ total_90✓ competitor_top✓ category_counts_present✓ side_split✓ two_patterns

✓ Correct Answer (ground truth)

Plain-English answer: Every closed-lost deal gets exactly one category (pricing, competitor, no decision, timing, product gap, champion left, other) plus a side, per the computed classification.
The answer in one line: total_lost_6mo = 90; top_reason = competitor (26)
total_lost_6mo90
categoriespricing, competitor, no decision, timing, product gap, champion left, other
patternstwo most actionable, tag-vs-text disagreements counted
sideBonusly / buyer / unknown per deal
top_reasoncompetitor (26)
graded checkstotal_90 · competitor_top · category_counts_present · side_split · two_patterns

Model's actual answer

show the model's full answer
Total deals reviewed: 90

Classification - alias | primary category | side
Deal-DB0AAC | timing | buyer
Deal-F7F635 | competitor | buyer
Deal-AC944F | no decision | unknown
Deal-214060 | no decision | unknown
Deal-91A056 | timing | buyer
Deal-29326C | timing | buyer
Deal-5DB9B0 | other | unknown - Spam / Does not fit ICP
Deal-831B7B | timing | buyer
Deal-F97C37 | competitor | buyer - other vendor more diversified
Deal-13E9CF | no decision | buyer - deprioritized, explicitly not budget
Deal-39E25C | timing | buyer
Deal-7ED004 | pricing | buyer - no budget approval
Deal-21B045 | no decision | unknown - MIA
Deal-B3ABED | timing | buyer - revisit Q2 next year / budget for 2028
Deal-422BA6 | competitor | buyer - ADP TotalSource preferred partner
Deal-ED9AE7 | other | buyer - Timing, budget, authority / Lost DM
Deal-988493 | no decision | unknown - mia
Deal-381C8C | competitor | buyer
Deal-F308CA | no decision | unknown - no contact since April
Deal-F1E8A6 | competitor | buyer
Deal-B6AC09 | timing | buyer - revisiting 2027
Deal-70F704 | no decision | unknown - anniversary-only + MIA
Deal-E6E80A | timing | buyer - pushed early 2027
Deal-B038F0 | timing | buyer - pushed early 2027
Deal-4664E1 | no decision | unknown - no contact after intro
Deal-175756 | timing | buyer - hold until 2027, other priorities
Deal-E74A73 | no decision | buyer - test manually before investing
Deal-DDAB52 | competitor | buyer - Rippl, more at same cost
Deal-ACE061 | competitor | buyer - feel HeyTaco
Deal-BB78F3 | timing | buyer - plant action items first
Deal-D48E0B | no decision | unknown - MIA
Deal-15DA99 | timing | buyer - early 2027
Deal-F4AF5D | timing | buyer - early next year
Deal-79B7A1 | timing | buyer
Deal-583ADB | no decision | unknown - MIA
Deal-8E27DA | no decision | buyer - swag only, didn't want R&R
Deal-2D2F8D | competitor | buyer
Deal-E0441F | no decision | unknown - stale, no contact
Deal-7CB44D | no decision | unknown - no contact since demo
Deal-0F96AA | competitor | buyer - RFP cut before finalist demo
Deal-1BCA50 | competitor | buyer - budget + gift cards, other stakeholder far with other vendor
Deal-7CC678 | competitor | buyer - no detail, tag only
Deal-FAC17C | other | buyer - no Exec IT Director approval
Deal-242273 | competitor | Bonusly - lost on digitize points currency + onsite spend differentiator
Deal-50E5D8 | no decision | buyer - pause
Deal-A2C349 | competitor | buyer - stick with Awardco + surveys
Deal-9F176A | timing | buyer - pause to end of year
Deal-7B2236 | pricing | buyer - simpler and cheaper
Deal-AFA56C | no decision | unknown - unresponsive
Deal-C7156E | competitor | buyer - selected another vendor
Deal-C33D91 | pricing | buyer - budget cuts
Deal-9048EB | product gap | Bonusly - bad fit + multiple feature gaps, no contact since April
Deal-5E64CE | competitor | buyer - locked in Nectar to Oct 2027
Deal-8A0992 | competitor | buyer - Canadian provider
Deal-D0C698 | competitor | buyer - wants Kudos again
Deal-69CF3D | timing | buyer - On Hold
Deal-ECBF89 | timing | buyer - On Hold
Deal-3618CC | product gap | Bonusly - Wanted Surveys
Deal-EECC02 | competitor | buyer
Deal-5AD03E | product gap | Bonusly - wanted more defined budget access
Deal-D1A623 | timing | buyer
Deal-413C56 | no decision | buyer - back to school priority, CEO not ready
Deal-47F1A1 | competitor | buyer - staying with WorkTango 12 mo
Deal-BF2A98 | competitor | buyer - deployed HiThrive
Deal-2A292B | no decision | buyer - build simple internally
Deal-D1AABF | no decision | unknown - No response
Deal-FEDBCB | no decision | buyer - not engaged, reopen if change
Deal-1E7DA9 | competitor | buyer - selected another platform
Deal-2BBA21 | no decision | unknown - no contact since intro
Deal-286F9C | competitor | buyer - another platform, not good fit
Deal-7FBAC6 | no decision | buyer - Leadership pause again
Deal-369281 | competitor | buyer - went with Paylocity
Deal-386F6E | no decision | unknown - No response
Deal-9FCD0D | competitor | buyer - Canadian company, CEO preference
Deal-55867E | no decision | buyer - not moving forward at this time, no future date given
Deal-DAFB82 | pricing | buyer - not budgeted until 2028, other priorities
Deal-2FEDDB | timing | buyer - unsure on timing to get moving
Deal-64B19A | competitor | buyer - likely stayed Motivosity
Deal-3F86A0 | no decision | unknown - unresponsive
Deal-096750 | no decision | unknown - no contact after intro
Deal-F325A5 | champion left | buyer - Layoffs and Change in Leadership
Deal-ABD14C | no decision | buyer - not interested
Deal-79E61A | no decision | unknown - Unresponsive
Deal-8A119B | pricing | buyer - Didn't get approval
Deal-AE7C4E | no decision | unknown - Unresponsive
Deal-DAB4F1 | no decision | unknown - Unresponsive
Deal-B4B50F | no decision | unknown - Unresponsive
Deal-981AD4 | product gap | Bonusly - Doesn't fit UI and not UK focused
Deal-DC77FE | competitor | Bonusly - more customization, label points as dollars, price not factor
Deal-5885B9 | no decision | unknown - MIA

SUMMARY

Category counts - arithmetic:
timing 19 + competitor 26 + no decision 32 + pricing 5 + product gap 4 + champion left 1 + other 3 = 90
timing: 19
competitor: 26
no decision: 32
pricing: 5
product gap: 4
champion left: 1
other: 3

Side split - arithmetic:
unknown 23 + Bonusly 6 + buyer 61 = 90
buyer: 61 = 19 timing + 24 competitor-buyer + 10 no-decision-buyer + 5 pricing + 1 champion + 2 other-buyer
Bonusly: 6 = Deal-9048EB, Deal-3618CC, Deal-5AD03E, Deal-981AD4 + Deal-242273, Deal-DC77FE
unknown: 23 = 22 no-decision MIA/unresponsive + Deal-5DB9B0 spam

Tag vs text disagreement: 8 deals where tag-implied category clearly differs from text:
Deal-70F704 - Lost DM vs no decision/MIA anniversary-only
Deal-8E27DA - Feature Request vs no decision, didn't want R&R
Deal-9048EB - MIA vs product gap/bad fit + feature gaps
Deal-5E64CE - Doing nothing vs competitor lock Nectar to 2027
Deal-3618CC - Lost DM vs product gap Wanted Surveys
Deal-5AD03E - Competitor vs product gap budget access, no competitor named
Deal-55867E - Timing vs no decision, no future date
Deal-2FEDDB - Doing nothing vs timing unsure

Two patterns to act on:
1. 2027 timing backlog: 19 timing losses, most explicitly early 2027 / 2028 incl Deal-91A056, Deal-B6AC09, Deal-E6E80A, Deal-B038F0, Deal-175756, Deal-15DA99, Deal-DAFB82. Needs dated nurture, not generic close-lost.
2. Early ghosting: 22 unknown-side no-decision, mostly no contact after intro/demo incl Deal-F308CA, Deal-4664E1, Deal-7CB44D, Deal-2BBA21, Deal-096750. Points to qualification / intro-to-next-step leak.
pipeline-tieringPASS 1.00
Plain-English question: Score all 156 open deals into exactly one of LOCK / ACTION / BUILD / REVIVE / WATCH / RISKY.
Prompt — what the model saw
Score the open pipeline below on these signals: forecast category, stage, meetings_30d, contact count, and engagement recency. Assign exactly one tier per deal: LOCK, ACTION, BUILD, REVIVE, WATCH, or RISKY (RISKY = the forecast category disagrees with the engagement evidence). Never tier a deal LOCK if it has zero meetings_30d. Do NOT list every deal. Return ONLY a compact JSON object with exactly these keys: - "tier_counts": object mapping each tier to the number of deals in it (must sum to the total deal count), - "tier_examples": object mapping each tier to up to 3 example deal aliases from the data, - "risky_deals": array of the deal aliases you tiered RISKY, - "lock_violations": number of deals tiered LOCK that have zero meetings_30d (should be 0), - "pipeline_shape": one short paragraph on the overall pipeline shape. Note: inbound_emails_30d is 0 for every row due to a data defect, so treat meetings_30d as the inbound signal.
cost $0.0222 · 424s · in 28,472 / out 35,112 tok
✓ json_returned✓ tier_vocabulary✓ tier_counts_sum✓ risky_deals_listed✓ no_lock_violation_claimed fabricated: Deal-C26D0

✓ Correct Answer (ground truth)

Plain-English answer: All 156 open deals tiered exactly once across the six tiers using forecast category, stage, meetings, contacts, and recency.
The answer in one line: total = 156
total156
graded checksjson_returned · tier_vocabulary · tier_counts_sum · risky_deals_listed · no_lock_violation_claimed

Model's actual answer

show the model's full answer
```json
{
  "tier_counts": {"LOCK": 3, "ACTION": 13, "BUILD": 45, "REVIVE": 32, "WATCH": 57, "RISKY": 6},
  "tier_examples": {
    "LOCK": ["Deal-D348E1", "Deal-C26D0", "Deal-403845"],
    "ACTION": ["Deal-25F752", "Deal-E53952", "Deal-C6D97A"],
    "BUILD": ["Deal-A5E80A", "Deal-C6FE92", "Deal-D73B89"],
    "REVIVE": ["Deal-7BBDFA", "Deal-42F601", "Deal-278DEC"],
    "WATCH": ["Deal-2D1F1B", "Deal-66D1FC", "Deal-950043"],
    "RISKY": ["Deal-547B2B", "Deal-B7EBD1", "Deal-A2B47C"]
  },
  "risky_deals": ["Deal-547B2B", "Deal-B7EBD1", "Deal-A2B47C", "Deal-2465CE", "Deal-584EE5", "Deal-FD9F4E"],
  "lock_violations": 0,
  "pipeline_shape": "156 open deals, $2,314,043 total. Only 16 deals ($147k, ~6%) are commit-closeable (LOCK+ACTION) and another 6 commit deals ($37.9k) are RISKY with zero meetings_30d — Deal-547B2B, Deal-B7EBD1, Deal-A2B47C, Deal-2465CE, Deal-584EE5, Deal-FD9F4E are forecast-committed but show outbound email only (e.g. Deal-A2B47C: 8 emails_30d, 0 meetings; last meeting 2026-07-27). The weight is in the middle/back: 45 BUILD deals ($520.6k) have fresh meeting evidence but sit in DS1–DS3, 32 REVIVE deals ($488.7k) are stale by stage age or recency (e.g. Deal-278DEC in DS2 for 211 days, Deal-42F601 for 224 days), and 57 WATCH deals ($1.12M, ~48% of value) are mostly DS1 — the three largest single bookings ($240k Deal-2D1F1B, $99k Deal-66D1FC, $70k Deal-950043) are all DS1 with zero meetings_30d. Net: thin near-term commit coverage against a large, unqualified top — revenue risk is forecast optimism at the bottom and an unworked bulk at the top."
}
```

Arithmetic check: 3+13+45+32+57+6 = 156 = total rows. LOCK requires ≥1 meetings_30d so lock_violations = 0 by construction. One correction to the alias above: LOCK examples are Deal-D348E1, Deal-C26D20, Deal-403845. Data defect noted: inbound_emails_30d = 0 for all rows (treated meetings_30d as the inbound signal); Deal-3EED2C and Deal-57FF13 have no engagement row at all, so they were scored on stage/forecast fields only.
call-transcript-extractionPASS 1.00
Plain-English question: Extract the CRM write-back fields from each call transcript as JSON.
Prompt — what the model saw
For each transcript, extract CRM write-back fields as JSON: why-buys (prospect statements only), pain points, stakeholders from the speaker list, budget signal (prospect-stated or null), timeline signal, competitor mentioned (only if the prospect raised it), next step (only if explicitly agreed), objections, confidence. Never fill a field from what the rep said or hoped.
cost $0.0024 · 66s · in 1 / out 4,449 tok
✓ json_emitted✓ budget_signal✓ null_discipline✓ stakeholders✓ competitor_prospect_only

✓ Correct Answer (ground truth)

Plain-English answer: Per transcript: prospect-only why-buys, pain points, stakeholders, budget and timeline signals, and competitors — e.g. TX-001: '$40k earmarked', live before open enrollment.
The answer in one line: rep = Alex Franklin
allowed_newTX-001, TX-002, TX-003, TX-004, TX-005, TX-006, Achievers, Bucketlist, Nectar
per_transcript{"TX-001": {"why_buys": ["automating anniversary and birthday awards"], "pain_points": ["HR team of three cannot keep up with it manually", "track everything in a spreadsheet, and people slip through
repAlex Franklin
rep_sourced_nulls{"TX-003": ["budget_signal \u2014 only the rep mentioned pricing/budget"], "TX-004": ["next_step \u2014 rep proposed, prospect said 'no promises'"], "TX-006": ["budget_signal \u2014 rep decoy, prospec
graded checksjson_emitted · budget_signal · null_discipline · stakeholders · competitor_prospect_only

Model's actual answer

show the model's full answer
[
  {
    "transcript_id": "TX-001",
    "deal_alias": "Deal-CFE7F4",
    "why_buys": [
      "Prospect (VP People): \"The big win for us would be automating anniversary and birthday awards — our HR team of three cannot keep up with it manually.\""
    ],
    "pain_points": [
      "HR team of three cannot keep up with anniversary/birthday awards manually",
      "Prospect (HR Admin): \"Right now we track everything in a spreadsheet, and people slip through the cracks.\""
    ],
    "stakeholders": ["Prospect (VP People)", "Prospect (HR Admin)"],
    "budget_signal": "Prospect (VP People): \"We have about $40k earmarked for engagement tools this fiscal year.\" (~$40,000)",
    "timeline_signal": "Prospect (VP People): live before open enrollment in November",
    "competitor_mentioned": {
      "name": "Achievers",
      "raised_by": "Prospect (VP People)",
      "note": "looked at it last year, too heavy for team size"
    },
    "next_step": "Security review with prospect's IT lead on September 12 — explicitly agreed by Prospect (VP People)",
    "objections": [
      "Prospect (HR Admin): need SSO and audit logs for IT to sign off"
    ],
    "confidence": "high"
  },
  {
    "transcript_id": "TX-002",
    "deal_alias": "Deal-70BB30",
    "why_buys": [
      "Prospect (Head of Total Rewards): \"We want to tie recognition to retention for our hourly workforce — regretted turnover there is over 30%.\""
    ],
    "pain_points": [
      "Regretted turnover over 30% in hourly workforce",
      "Recognition not tied to retention"
    ],
    "stakeholders": ["Prospect (Head of Total Rewards)", "Prospect (CFO)"],
    "budget_signal": "Prospect (CFO): \"Finance has approved a $25k pilot budget for this quarter.\" ($25,000, approved)",
    "timeline_signal": "Prospect (CFO): decision by end of September",
    "competitor_mentioned": null,
    "next_step": "Rep to send pilot agreement; Prospect (CFO) agreed: \"send the pilot agreement and we'll route it to legal this week.\"",
    "objections": [
      "Prospect (CFO): \"Integration with Workday has to be rock solid — that's my one condition.\""
    ],
    "confidence": "high"
  },
  {
    "transcript_id": "TX-003",
    "deal_alias": "Deal-530B50",
    "why_buys": [
      "Prospect (People Ops Manager): \"We need to make recognition visible across our 12 retail locations.\"",
      "Prospect (People Ops Manager): store managers need budget autonomy for on-the-spot recognition"
    ],
    "pain_points": [
      "Recognition not visible across 12 retail locations",
      "Store managers have zero budget autonomy for on-the-spot recognition today"
    ],
    "stakeholders": ["Prospect (People Ops Manager)"],
    "budget_signal": null,
    "timeline_signal": "Prospect (People Ops Manager): \"Honestly there's no rush on our side until Q1.\"",
    "competitor_mentioned": {
      "name": "Bucketlist",
      "raised_by": "Prospect (People Ops Manager)",
      "note": "CEO used it at her last company and liked it"
    },
    "next_step": "Schedule a call with the prospect's CEO — Prospect (People Ops Manager) agreed to send two times",
    "objections": [
      "No urgency until Q1",
      "Prospect (People Ops Manager): \"The CEO has to be sold first — she decides anything people-related.\""
    ],
    "confidence": "medium"
  },
  {
    "transcript_id": "TX-004",
    "deal_alias": "Deal-180D02",
    "why_buys": [
      "Prospect (VP People): \"We want to consolidate three separate recognition tools into one.\""
    ],
    "pain_points": [
      "Paying for three tools and none of them talk to their HRIS",
      "Prior security review took three months"
    ],
    "stakeholders": ["Prospect (VP People)", "Prospect (IT Security Lead)"],
    "budget_signal": "Prospect (VP People): \"If it's under $15k annually, I can approve it without going to the board.\" (threshold, not a committed budget)",
    "timeline_signal": "Prospect (IT Security Lead): procurement cycle runs six to eight weeks minimum",
    "competitor_mentioned": null,
    "next_step": null,
    "objections": [
      "Prospect (IT Security Lead): security review took three months for last vendor — stated hesitation",
      "Prospect (VP People) declined to commit to CFO follow-up: \"Maybe — I need to check her calendar, no promises.\""
    ],
    "confidence": "medium"
  },
  {
    "transcript_id": "TX-005",
    "deal_alias": "Deal-F8767A",
    "why_buys": [
      "Prospect (HR Director): \"Two things: automate service milestones, and give us analytics on recognition equity across departments.\""
    ],
    "pain_points": [
      "Service milestones not automated",
      "Recognition equity gaps across departments",
      "Prospect (People Ops Coordinator): night-shift teams feel invisible — engagement scores run 20 points lower"
    ],
    "stakeholders": ["Prospect (HR Director)", "Prospect (People Ops Coordinator)"],
    "budget_signal": "Prospect (HR Director): \"We have $12k approved under our engagement line.\" ($12,000, approved)",
    "timeline_signal": "Prospect (HR Director): running before the January all-hands",
    "competitor_mentioned": {
      "name": "Nectar",
      "raised_by": "Prospect (HR Director)",
      "note": "mid-pilot with Nectar; \"you'd need to beat that experience\""
    },
    "next_step": "Rep to present directly to prospect's exec team on October 2 — explicitly agreed by Prospect (HR Director)",
    "objections": [
      "Prospect (HR Director): exec team skeptical after a failed rollout two years ago",
      "Incumbent Nectar pilot to beat"
    ],
    "confidence": "high"
  },
  {
    "transcript_id": "TX-006",
    "deal_alias": "Deal-EE195F",
    "why_buys": [
      "Prospect (HR Manager): \"I'd love to cut the admin time on service awards.\""
    ],
    "pain_points": [
      "Prospect (HR Manager) personally spends five hours a month ordering and shipping plaques"
    ],
    "stakeholders": ["Prospect (HR Manager)"],
    "budget_signal": "Prospect (HR Manager): \"Budget isn't the issue — time is.\" (no amount stated)",
    "timeline_signal": "Prospect (HR Manager): fiscal year resets in February, Q1 start is realistic",
    "competitor_mentioned": null,
    "next_step": "Rep to send one-page overview; Prospect (HR Manager) agreed to forward it to the COO this week",
    "objections": [
      "Prospect (HR Manager): \"our COO usually prefers building things in-house\" — comparing against doing it internally"
    ],
    "confidence": "medium"
  }
]

Extraction notes applied uniformly: budget rep-quoted pricing (TX-003, "$8 per employee per month") excluded as rep-stated; TX-004 next step is null because the rep's proposed CFO call was met with "no promises" — not an explicit agreement; TX-006's build-vs-buy stance is logged as an objection, not a competitor, since no vendor was named.
next-to-closePASS 1.00
Plain-English question: Which 3 deals are closest to signature right now?
Prompt — what the model saw
Which 3 deals are closest to signature right now? Use the CRM data and the Slack thread; Slack often has the real status before the CRM does. For each: deal, amount, why it is close, what is left. Under 250 words.
cost $0.0013 · 30s · in 1,731 / out 1,636 tok
✓ exactly_three✓ top3_named✓ slack_first✓ blocked_excluded

✓ Correct Answer (ground truth)

Plain-English answer: Deal-547B2B ($11,200), Deal-B7EBD1 ($9,000), and Deal-403845 ($9,000) — all DS5 COMMIT deals closing September 10-11, with what is left on each.
blocked_excludeddeal_alias=Deal-2465CE; reason=champion left + procurement freeze (Slack #sales-sparring 2026-09-04)
exactly_three1
ranking_ruleSlack signature-imminent deal ranks first; remaining ranked by stage (DS5>DS4), forecast (COMMIT>BEST_CASE>PIPELINE), close_date asc, amount desc
top3rank=1; deal_alias=Deal-547B2B; amount=11200; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11 | rank=2; deal_alias=Deal-B7EBD1; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-10 | rank=3; deal_alias=Deal-403845; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11
graded checksexactly_three · top3_named · slack_first · blocked_excluded

Model's actual answer

show the model's full answer
**Three closest to signature** (DS5/COMMIT, nearest close dates, Slack-confirmed where possible):

1. **Deal-547B2B** — $11,200 (Alex Franklin, DS5/COMMIT, close 2026-09-11)
   Why close: Slack 09-04 — "redlines came back clean… signing page is out… they are signing tomorrow." Only deal called "signature-imminent."
   What's left: Nothing identified beyond countersignature — no open blocker cited.

2. **Deal-403845** — $9,000 (Alex Franklin per CRM; Dana reports status, DS5/COMMIT, close 2026-09-11)
   Why close: DS5/COMMIT + Slack: "moving fine… order form is with their finance team."
   What's left: Finance sign-off on the order form, then signature.

3. **Deal-B7EBD1** — $9,000 (Dana Mercer, DS5/COMMIT, close 2026-09-10)
   Why close: Earliest close date in the file, DS5/COMMIT, no negative signal in Slack. Weakest of the three — no Slack corroboration at all.
   What's left: Unknown from provided data.

Sum of the three: $11,200 + $9,000 + $9,000 = **$29,200**.

**Excluded despite CRM rank:** Deal-2465CE ($5,400, DS5/COMMIT 09-10) — Slack overrides the CRM: champion left, procurement froze, Dana is "pulling it out of commit… realistically a Q4 deal now." Deal-A2B47C ($6,360) is "still warm, just normal legal-review pace" but trails on amount and has no near-term signing signal.

**Data caveats:** The Slack export is dated 2026-09-04 and today is 2026-09-25 — all three picks carry CRM close dates that have already passed with no later status shown. Whether they actually signed is missing from the provided data; the CRM file may simply be stale for these records.
product-gap-detectionPASS 1.00
Plain-English question: Which prospect complaints are true product gaps versus rollout or enablement issues?
Prompt — what the model saw
Surface true product gaps raised by prospects in active deals. For each candidate: quote the prospect line, classify as real gap, plan gate, or rollout/enablement issue using the product docs excerpt, and name the deal and amount. Only real gaps go in the summary. Rep language does not count as prospect voice.
cost $0.0013 · 33s · in 1,001 / out 1,746 tok
✓ real_gaps✓ plan_gate_classified✓ enablement_classified✓ gap_summary

✓ Correct Answer (ground truth)

Plain-English answer: Two real gaps (SCIM provisioning, ADP Workforce Now integration) separated from plan gates and enablement issues, each with the prospect quote and deal.
allowed_newTG-001, TG-002, TG-003, TG-004, TG-005, SCIM user provisioning, ADP Workforce Now integration, Custom report builder
enablement_issuestranscript_id=TG-004; deal_alias=Deal-84DBA6; amount=16000; feature=Slack integration recognition adoption; quote=We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.; why_enablement=Slack integration exists on all plans; managers were never trained
excluded_rep_voicetranscript_id=TG-005; line=Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap.; reason=rep voice, not prospect-raised
plan_gatestranscript_id=TG-003; deal_alias=Deal-CFE7F4; amount=32000; feature=Custom report builder; tier_required=Enterprise; quote=I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?
real_gapstranscript_id=TG-001; deal_alias=Deal-EC3025; amount=62000; feature=SCIM user provisioning; quote=We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.; why_real=not listed in product docs at any tier | transcript_id=TG-002; deal_alias=Deal-D0D6B5; amount=23800; feature=ADP Workforce Now integration; quote=Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.; why_real=supported HRIS list is Workday/BambooHR/Gusto/Rippling only
graded checksreal_gaps · plan_gate_classified · enablement_classified · gap_summary

Model's actual answer

show the model's full answer
GT classification complete. One data limitation up front: deal amounts are not present in either provided file — only deal aliases. I flag the amount as MISSING for every candidate rather than inventing it.

CANDIDATE-BY-CANDIDATE

TG-001 / Deal-EC3025 (amount: not in provided data)
Quote: "We need SCIM user provisioning — our security team will not approve any tool that requires manual account management."
Classification: REAL GAP. The docs excerpt (snapshot 2026-08) contains no mention of SCIM or automated user provisioning anywhere, and the prospect frames it as a security-approval blocker. Caveat stated honestly: the excerpt covers only plans/tiers, so absence of SCIM is absence-of-evidence in the given docs, not a confirmed "not supported" statement.

TG-002 / Deal-D0D6B5 (amount: not in provided data)
Quote: "Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us."
Classification: REAL GAP. The docs excerpt includes no integration list at all, and no ADP reference exists in it; the prospect's own statement places ADP outside the published integration list. Same caveat: cannot verify the full integration list from the data provided.

TG-003 / Deal-CFE7F4 (amount: not in provided data)
Quote: "I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?"
Classification: PLAN GATE. The capability exists (rep identifies the custom report builder as an Enterprise-tier feature; the docs table structure confirms tiered capability gating); the prospect is blocked by plan, not by a missing product. Note: the Enterprise-tier placement itself comes from rep language, and the docs excerpt's tier table shows only one capability row (peer-to-peer recognition + points catalog), so tier placement is not independently verifiable from the excerpt.

TG-004 / Deal-84DBA6 (amount: not in provided data)
Quote: "We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it."
Classification: ROLLOUT/ENABLEMENT ISSUE. The Slack integration is working and already turned on; the failure is adoption/training, not product.

TG-005 / Deal-36C33F (amount: not in provided data)
Quote (prospect only): "Good to know. The web version should be fine for our office staff for now."
Classification: EXCLUDED. The gap (no native mobile app) was raised by the rep, not the prospect; the prospect explicitly said it is not blocking. Rep language does not count as prospect voice.

TALLY (arithmetic)
Candidates surfaced: 5. Excluded: 1 (TG-005). Remaining: 5 − 1 = 4. Split: 2 real gaps + 1 plan gate + 1 enablement = 4. ✓ Consistent.

SUMMARY — TRUE PRODUCT GAPS (prospect-raised only)
1. SCIM user provisioning — Deal-EC3025, amount MISSING from provided data. Security-team approval blocker.
2. ADP Workforce Now integration — Deal-D0D6B5, amount MISSING from provided data. Prospect-declared dealbreaker.

To attach amounts and verify the two gaps against the full (non-excerpted) integration/feature docs, that data would need to be supplied — e.g., a HubSpot deal lookup on the two aliases.
rep-performance 5 tests
stale-pipeline-by-rep0.83
Plain-English question: Which open deals have had no email, call, or meeting in the last 7 days?
Prompt — what the model saw
The snapshot date is 2026-09-05. Using the open deals and the engagements-by-deal table, list every open deal with no logged email, call, or meeting in the last 7 days. Compute recency from the engagements table's last_email/last_call/last_meeting fields, NOT from any date field on the deal (the deal's last_contacted_field is unreliable). For each stale deal give: deal alias, owner name, stage, amount, days since last contact. Group by owner, order by amount descending within each owner. Then state, per owner, how many stale deals they carry and the total stale amount. (Do not write outreach text.)
cost $0.0340 · 600s · in 31,787 / out 56,352 tok
✓ stale_recall_ge_80✓ grouped_by_owner✓ jake_top_amount✓ kevin_stale_amt✓ jenn_stale_amt✗ stale_count

✓ Correct Answer (ground truth)

Plain-English answer: 65 open deals are stale; the largest is Deal-2D1F1B at $240,000, with Bryce Harmon carrying the most stale dollars at $626,243.
The answer in one line: jake_amt = 626,243; jenn_amt = 261,645; kevin_amt = 252,905.03; top_amount = 240,000; stale_count = 65; top_alias = Deal-2D1F1B
jake_amt626,243
jenn_amt261,645
kevin_amt252,905.03
stale_count65
top_amount240,000
top_aliasDeal-2D1F1B
graded checksstale_recall_ge_80 · grouped_by_owner · jake_top_amount · kevin_stale_amt · jenn_stale_amt · stale_count

Model's actual answer

show the model's full answer
Method: snapshot 2026-09-05; recency = max(last_email, last_call, last_meeting) from engagements_by_deal_90d.csv (deal last_contacted_field ignored). A deal is stale if that max ≤ 2026-08-28 (i.e., ≥8 days before snapshot; the "last 7 days" window = 08-29 → 09-05). Days = 2026-09-05 minus max date. No deal's max lands exactly on 08-29, so the boundary choice doesn't affect the list.

Missing data (stated explicitly): Deal-3EED2C (Alex Franklin, DS2, $7,200) and Deal-57FF13 (Elena Sinclair, DS1, $2,100) have NO row in the engagements table — recency is unverifiable, so they are EXCLUDED from the stale lists and totals below. Elena Sinclair has no engagements row at all; Farid Osman has zero stale deals.

BRYCE HARMON — 13 stale, $626,243
- Deal-2D1F1B | DS1 | $240,000 | 81d (last 2026-06-16: email 06-11 < meeting 06-16)
- Deal-66D1FC | DS1 | $99,000 | 16d (08-20)
- Deal-950043 | DS1 | $70,000 | 19d (08-17)
- Deal-B23205 | DS1 | $45,000 | 16d (08-20)
- Deal-7BBDFA | DS3 | $37,440 | 46d (07-21)
- Deal-332637 | DS2 | $36,000 | 9d (08-27)
- Deal-1BEEBF | DS1 | $31,500 | 19d (08-17 email; call 07-30 older)
- Deal-C5658B | DS1 | $23,400 | 16d (08-20)
- Deal-40522D | DS3 | $21,000 | 19d (08-17)
- Deal-F0EBBB | DS3 | $11,400 | 24d (08-12)
- Deal-E25A09 | DS1 | $6,000 | 9d (08-27)
- Deal-C9C286 | DS2 | $5,502 | 9d (08-27)
- Deal-012CB1 | DS1 | $1 | 23d (08-13)
Sum: 240000+99000+70000+45000+37440+36000+31500+23400+21000+11400+6000+5502+1 = $626,243

ALEX FRANKLIN — 18 stale, $102,336 (+1 unverifiable: Deal-3EED2C, $7,200)
- Deal-CC08D1 | DS1 | $24,000 | 16d (08-20)
- Deal-E73427 | DS3 | $18,000 | 10d (08-26)
- Deal-885F45 | DS2 | $9,300 | 12d (08-24)
- Deal-C2FF3C | DS1 | $8,316 | 10d (08-26)
- Deal-0D2F7A | DS3 | $5,100 | 12d (08-24 call > 08-05 email)
- Deal-6C60D4 | DS3 | $4,800 | 12d (08-24 call > 07-31 email)
- Deal-13FEBD | DS2 | $4,680 | 12d (08-24 call > 08-04 email)
- Deal-9D0060 | DS3 | $3,840 | 12d (08-24)
- Deal-690476 | DS2 | $3,600 | 18d (08-18 call > 08-03 email)
- Deal-C6D97A | DS4 | $3,240 | 8d (08-28 email > 08-25 call)
- Deal-EE195F | DS3 | $3,120 | 8d (08-28)
- Deal-278DEC | DS3 | $2,700 | 8d (08-28)
- Deal-635B8E | DS3 | $2,600 | 18d (08-18)
- Deal-6883F3 | DS1 | $2,400 | 16d (08-20)
- Deal-4A13AD | DS3 | $2,160 | 26d (08-10)
- Deal-F67D31 | DS2 | $1,800 | 8d (08-28)
- Deal-5FDCE4 | DS3 | $1,600 | 12d (08-24)
- Deal-BA571A | DS4 | $1,080 | 18d (08-18)
Sum = $102,336

DANA MERCER — 14 stale, $261,645
- Deal-44EA29 | DS2 | $60,000 | 10d (08-26)
- Deal-E51FB7 | DS2 | $43,875 | 12d (08-24 call > 08-18 email)
- Deal-B42F46 | DS1 | $27,000 | 19d (08-17)
- Deal-BA3DDC | DS3 | $23,400 | 15d (08-21 call > 08-20 email)
- Deal-9DDE86 | DS2 | $20,000 | 15d (08-21)
- Deal-215CCA | DS3 | $18,900 | 17d (08-19 meeting; emails only to 07-02)
- Deal-5EED42 | DS3 | $16,250 | 11d (08-25)
- Deal-57887A | DS2 | $15,000 | 8d (08-28)
- Deal-B7EBD1 | DS5 | $9,000 | 16d (08-20)
- Deal-3974EB | DS4 | $9,000 | 8d (08-28)
- Deal-F40F04 | DS2 | $8,100 | 15d (08-21)
- Deal-87DDD1 | DS1 | $5,000 | 19d (08-17)
- Deal-F336B6 | DS3 | $4,200 | 15d (08-21 email > 08-18 meeting)
- Deal-0660B4 | DS4 | $1,920 | 26d (08-10 email > 08-20?? no: last_meeting 08-20; last_email 08-10 → actual max 08-20 = 16d)
  Correction: Deal-0660B4 | DS4 | $1,920 | 16d (08-20 meeting > 08-10 email)
Sum: 60000+43875+27000+23400+20000+18900+16250+15000+9000+9000+8100+5000+4200+1920 = $261,645

COLE INGRAM — 18 stale, $252,905.03
- Deal-D04904 | DS2 | $58,529.25 | 11d (08-25)
- Deal-B25F40 | DS3 | $40,000 | 8d (08-28)
- Deal-813836 | DS2 | $32,175 | 11d (08-25)
- Deal-1BA595 | DS2 | $31,750 | 11d (08-25)
- Deal-CFE1E8 | DS3 | $18,000 | 11d (08-25)
- Deal-CD47A6 | DS2 | $12,168 | 11d (08-25 call; email 08-25)
- Deal-627646 | DS3 | $11,193 | 11d (08-25)
- Deal-FF809F | DS2 | $7,781.20 | 11d (08-25)
- Deal-AF932D | DS2 | $7,225.40 | 11d (08-25)
- Deal-A71728 | DS2 | $6,947.50 | 11d (08-25)
- Deal-8BC9F5 | DS2 | $5,616 | 10d (08-26)
- Deal-175395 | DS3 | $4,779.88 | 11d (08-25)
- Deal-481E24 | DS3 | $4,140 | 10d (08-26 call > 08-25 email)
- Deal-C7F9BF | DS2 | $3,360 | 11d (08-25 email > 08-24 call)
- Deal-2F3A66 | DS3 | $3,334.80 | 11d (08-25)
- Deal-342E96 | DS2 | $2,700 | 24d (08-12)
- Deal-E568D5 | DS3 | $1,875 | 11d (08-25)
- Deal-FD9F4E | DS5 | $1,330 | 10d (08-26)
Sum = $252,905.03

FARID OSMAN — 0 stale, $0 (all 7 of his open deals show ≥1 email/call/meeting on or after 08-25/08-31 window edge or later; earliest qualifying contact is within 7 days via dated activity, e.g. Deal-499BF6 email 09-03, Deal-F17780 09-01, Deal-2B39B0 09-03; Deal-03BA69 email 08-25 + meeting 09-02)

ELENA SINCLAIR — 0 verifiable stale (sole deal Deal-57FF13 has no engagements row)

ROLL-UP (excluding the 2 unverifiable deals)
| Owner | Stale deals | Stale $ |
|---|---|---|
| Bryce Harmon | 13 | $626,243.00 |
| Alex Franklin | 18 | $102,336.00 (+$7,200 unverifiable) |
| Dana Mercer | 14 | $261,645.00 |
| Cole Ingram | 18 | $252,905.03 |
| Farid Osman | 0 | $0 |
| Elena Sinclair | 0 | ($2,100 unverifiable) |
| TOTAL | 63 | $1,243,129.03 |
Check: 626,243 + 102,336 + 261,645 + 252,905.03 = $1,243,129.03

Data-quality caveat: 17 deals have last_meeting dates AFTER the 2026-09-05 snapshot (e.g., Deal-944310 meeting 2026-09-15, Deal-3795AD 2026-10-02). Taken as given, those deals are not stale. If future-dated entries are instead discarded (likely scheduled meetings, not logged contact), 10 more deals flip to stale using their latest ≤09-05 activity: Deal-01E193 ($12,600, 8d), Deal-927338 ($10,920, 18d), Deal-A414F6 ($25,200, 19d), Deal-C1FA6D ($18,000, 16d), Deal-3795AD ($1, 8d) [Bryce → 18 stale, $692,964]; Deal-944310 ($10,500, 33d), Deal-7599B8 ($7,350, 18d) [Dana → 16, $279,495]; Deal-819506 ($4,400, 8d) [Alex → 19, $106,736]; Deal-8BA24E ($2,880, 8d), Deal-8FDCD2 ($1, 15d) [Farid → 2, $2,881]. Cole unchanged.
activity-mix-vs-outcomePASS 1.00
Plain-English question: Per rep, what did the last 30 days of emails, calls, and meetings look like next to the deals that advanced?
Prompt — what the model saw
The snapshot date is 2026-09-05. For each sales rep (owner), compute the last-30-day totals of emails, calls, and meetings from the engagements table, and the number of deals that entered DS2 in the last 30 days (t_ds2 within 30 days of the snapshot). Report per rep: the activity mix as percentages (emails/calls/meetings share of that rep's total activities), total activities per DS2 entry, and rank the reps by that efficiency ratio (lowest activities per DS2 entry = most efficient). Name the most efficient rep and the highest-volume rep, and note if they differ. Do not attribute cause.
cost $0.0267 · 418s · in 34,050 / out 37,811 tok
✓ alex_ds2_30d✓ justin_volume_high✓ efficient_named✓ mix_present✓ no_causal_claim

✓ Correct Answer (ground truth)

Plain-English answer: Alex Franklin is the standout signal with 18 deals entering DS2 in 30 days (next closest: Bryce Harmon with 4), alongside per-rep activity totals.
The answer in one line: alex_ds2_30d = 18; bryce_ds2_30d = 4; note = legacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
alex_ds2_30d18
bryce_ds2_30d4
notelegacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
graded checksalex_ds2_30d · justin_volume_high · efficient_named · mix_present · no_causal_claim

Model's actual answer

show the model's full answer
All computed. One data-integrity note up front, then the results.

Window: last 30 days = t_ds2 in [2026-08-06, 2026-09-05] (snapshot 2026-09-05).

Data caveats (stated, not invented around):
- Deal-3EED2C (Alex Franklin) entered DS2 on 2026-09-03 and counts toward the DS2-entry denominator, but it has NO row in engagements_by_deal_30d — its activity is therefore absent from the numerator. (Excluding it entirely would give Alex 384/17 = 22.59; ranking is unchanged either way.)
- Deal-57FF13 (Elena Sinclair) also has no engagements row. Elena has 0 recorded activities and 0 DS2 entries → ratio undefined.

Per rep (emails_30d / calls_30d / meetings_30d summed over each rep's deals; inbound emails are 0 for every row):

Bryce Harmon
  E 162, C 0, M 43 → total 205
  Mix: 162/205 = 79.0% emails, 0/205 = 0.0% calls, 43/205 = 21.0% meetings
  DS2 entries (30d): 4 (Deal-25F752 08-10, Deal-CA7DC0 08-12, Deal-1CCE5C 08-06, Deal-D73B89 09-03)
  Activities per DS2 entry: 205/4 = 51.25

Alex Franklin
  E 307, C 36, M 41 → total 384
  Mix: 307/384 = 79.9% emails, 36/384 = 9.4% calls, 41/384 = 10.7% meetings
  DS2 entries (30d): 18 — Deal-EE195F, Deal-D9A72E (both 08-06), Deal-7FA0C3, Deal-E531A6 (08-07), Deal-D1E6C2, Deal-36C33F (08-11), Deal-317E6F (08-12), Deal-F436DA (08-19), Deal-4F775F (08-17), Deal-CA5E44 (08-24), Deal-46988D (08-26), Deal-5296C9, Deal-898FC5, Deal-E73427 (08-28), Deal-403845, Deal-92D97D (09-02), Deal-1FC049, Deal-3EED2C (09-03)
  Activities per DS2 entry: 384/18 = 21.33

Dana Mercer
  E 84, C 18, M 11 → total 113
  Mix: 84/113 = 74.3% emails, 18/113 = 15.9% calls, 11/113 = 9.7% meetings
  DS2 entries (30d): 1 (Deal-57887A 08-07)
  Activities per DS2 entry: 113/1 = 113.00

Cole Ingram
  E 96, C 14, M 1 → total 111
  Mix: 96/111 = 86.5% emails, 14/111 = 12.6% calls, 1/111 = 0.9% meetings
  DS2 entries (30d): 2 (Deal-1BA595 08-12, Deal-42326B 08-26)
  Activities per DS2 entry: 111/2 = 55.50

Farid Osman
  E 38, C 0, M 34 → total 72
  Mix: 38/72 = 52.8% emails, 0/72 = 0.0% calls, 34/72 = 47.2% meetings
  DS2 entries (30d): 1 (Deal-499BF6 08-26)
  Activities per DS2 entry: 72/1 = 72.00

Elena Sinclair
  E 0, C 0, M 0 → total 0 (only deal Deal-57FF13 has no engagements row)
  DS2 entries (30d): 0
  Ratio: undefined (no activities, no DS2 entries)

Ranking by activities per DS2 entry (lowest = most efficient):
  1. Alex Franklin — 21.33
  2. Bryce Harmon — 51.25
  3. Cole Ingram — 55.50
  4. Farid Osman — 72.00
  5. Dana Mercer — 113.00
  —. Elena Sinclair — n/a

Most efficient rep: Alex Franklin (21.33 activities per DS2 entry).
Highest-volume rep (total activities): Alex Franklin (384; next is Bryce Harmon at 205).
They are the SAME rep — Alex is both the highest-volume and (by this ratio) the most efficient, driven by 18 DS2 entries vs everyone else's ≤4. No causal attribution implied.
ae-qtd-scorecardPASS 1.00
Plain-English question: Build Alex Franklin's QTD scorecard: bookings versus quota, pipeline, and conversion.
Prompt — what the model saw
Build a QTD scorecard for the named AE as of 2026-09-05: bookings vs quota with attainment percent, new vs expansion split, active pipeline by stage with amount, rolling 90-day DS2-to-won rate, win and loss counts with top loss reason, activity volume by type in the last 30 days. Exclude any closed-won deal dated before the quarter. Close with three coaching observations grounded in the numbers.
cost $0.0472 · 309s · in 19,273 / out 27,247 tok
✓ bookings✓ quota✓ attainment✓ new_expansion_split✓ top_loss_reason✓ coaching_obs

✓ Correct Answer (ground truth)

Plain-English answer: 75% attainment ($150,000 on a $200,000 quota), $113,500 new versus $36,500 expansion, pipeline by stage, and the rolling DS2-to-won rate.
The answer in one line: quota = 200000; bookings = 150000; new_bookings = 113500; expansion_bookings = 36500; open_deals_count = 125; attainment_pct = 75; owner = Alex Franklin (+12 more below)
quota200,000
attainment_pct75
bookings150,000
expansion_bookings36,500
expansion_count3
losses_qtd27
new_bookings113,500
new_count5
open_deals_count125
wins_qtd8
activity_30d_totalsemails_30d=807; calls_30d=112; meetings_30d=128; notes_30d=50
allowed_new18 items — e.g. Deal-A1C3E5, Deal-B7D2F4, Deal-C9E1A6, Deal-D4B8C2
ds2_to_wonwindow_start=2026-06-07; won_in_window=8; closed_in_window=35; rate=0.2286
ownerAlex Franklin
pipeline_by_stage{"DS1": {"count": 20, "amount": 284621}, "DS2": {"count": 28, "amount": 353760}, "DS3": {"count": 67, "amount": 552705}, "DS4": {"count": 5, "amount": 23574}, "DS5": {"count": 5, "amount": 45730}}
prior_quarter_excludeddeal_alias=Deal-B3E6F1; close_date=2026-06-20; amount=24000
quarter2026-Q3
snapshot_date2026-09-05
top_loss_reasonreason=Lost- Timing (1 year or more); count=13
graded checksbookings · quota · attainment · new_expansion_split · top_loss_reason · coaching_obs

Model's actual answer

show the model's full answer
QTD Scorecard — Alex Franklin — as of 2026-09-05 — 2026-Q3

Bookings vs quota:
Quota $200,000. Bookings $150,000. Attainment 75.0%.
Arithmetic: 40,000 + 20,000 + 35,000 + 21,000 + 12,000 + 11,000 + 6,500 + 4,500 = 150,000. 150,000 / 200,000 = 75.0%.
Excluded pre-quarter won: Deal-B3E6F1, $24,000, closed 2026-06-20.

New vs expansion (won QTD only):
New: 5 deals, $113,500 — Deal-A1C3E5 $40,000 + Deal-B7D2F4 $35,000 + Deal-C9E1A6 $21,000 + Deal-D4B8C2 $11,000 + Deal-E6F3A9 $6,500 = $113,500.
Expansion: 3 deals, $36,500 — Deal-F2C7D8 $20,000 + Deal-A8B4D6 $12,000 + Deal-C5D9E2 $4,500 = $36,500.
deal_type is blank in file for all open/lost rows, so split available for won only.

Active pipeline (status = open, 125 deals, $1,260,390):
DS1: 20 deals, $284,621
DS2: 28 deals, $353,760
DS3: 67 deals, $552,705
DS4: 5 deals, $23,574
DS5: 5 deals, $45,730
Total: 284,621 + 353,760 + 552,705 + 23,574 + 45,730 = 1,260,390.

Rolling 90-day DS2-to-won rate:
Window: 2026-06-07 to 2026-09-05. Entered DS2 in window: 111 deals. Won from cohort: 8.
Rate: 8 / 111 = 7.2%.
Still open from cohort: 76. Lost from cohort: 27. Decided-only: 8 / 35 = 22.9%.

Wins and losses QTD (close_date 2026-07-01 to 2026-09-05):
Wins: 8. Losses: 27, $329,272.
Loss reasons: Lost- Timing (1 year or more) 13, MIA 5, Competitor 5, Lost DM 2, Feature Request 1, Lost- Does not fit ICP (write in notes) 1.
Top reason: Lost- Timing (1 year or more) — 13 / 27 = 48.1%, $184,681.

Activity last 30 days (sum ae_engagements.csv, 161 rows):
Emails 807, Calls 112, Meetings 128, Notes 50. Total 807 + 112 + 128 + 50 = 1,097.

Coaching observations:
1. Qualification: 13 of 27 losses are Lost- Timing (1 year or more), $184,681 lost — tighten DS2 entry timing criteria before building $552,705 DS3 bulk.
2. Conversion: 7.2% DS2-to-won with 76 of 111 cohort still open and 28 open DS2 at $353,760 — prioritize advancing/deciding aged DS2, not adding DS1 (20 deals, $284,621).
3. Mix: 807 emails vs 112 calls + 128 meetings — shift late-stage DS3/DS4/DS5 effort to live selling to lift 22.9% decided win rate.
multithreading-gapPASS 1.00
Plain-English question: Which open deals are single-threaded or under-threaded?
Prompt — what the model saw
Find every open deal that is single-threaded (fewer than 2 active contacts) or under-threaded (fewer than 3, or all contacts in one persona). Active means engaged in the last 60 days and not marked former. Personas: economic buyer, champion, HR admin, IT security, finance. For each flagged deal: amount, stage, active contact count, personas present, personas missing, the most valuable persona to add given the stage, and any on-file unengaged contact at that company who fits; otherwise say none on file.
cost $0.0039 · 113s · in 241 / out 7,390 tok
✓ flagged_recall✓ persona_language✓ single_thread_term✓ on_file_fit

✓ Correct Answer (ground truth)

Plain-English answer: Deals with fewer than 2 active contacts (single-threaded) or fewer than 3, or all in one persona (under-threaded), each flagged with the missing persona.
The answer in one line: active_window_start = 2026-07-07
active_window_start2026-07-07
allowed_new59 items — e.g. C-077A0E, C-0D15DF, C-1BDFE0, C-2AE3AA
best_persona_ruleDS2/DS3: economic buyer > IT security > finance > HR admin > champion; DS4/DS5: economic buyer > finance > IT security > HR admin > champion; first persona not present among active contacts
flag_ruleflag if active_contacts < 2 (single-threaded) or active_contacts < 3 or all active contacts share one persona (under-threaded); active = engaged within 60 days and not former
flaggeddeal_alias=Deal-36C33F; amount=15000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-50D386; amount=36000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5408B0; amount=14850; stage=DS4; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5BFE3B; amount=18000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-885F45; amount=9300; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=IT security | deal_alias=Deal-92D97D; amount=60000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer … (+5 more)
healthy_not_flaggedDeal-4B0BEB, Deal-84DBA6, Deal-D348E1
persona_universeeconomic buyer, champion, HR admin, IT security, finance
graded checksflagged_recall · persona_language · single_thread_term · on_file_fit

Model's actual answer

show the model's full answer
Ran the threading analysis on the two files. Two data limitations up front, stated explicitly:

- AMOUNT AND STAGE ARE NOT IN THE PROVIDED DATA. Neither CSV contains a deal amount or a deal stage field, so those columns read "not in data" for every deal, and "most valuable persona given the stage" cannot be stage-ranked. What the data does support: 10 of the 11 flagged deals have NO active economic buyer (all but Deal-885F45), so the economic buyer is the missing persona in every case except Deal-885F45, where the gap is HR admin / IT security / finance.
- No open/closed field was provided; all 14 deals in deal_contacts.csv were treated as in scope.

Arithmetic / method:
- Reference date = today, 2026-09-25. 60-day cutoff: 2026-09-25 − 60 days = 2026-07-27 (Sep 25 + Aug 31 = 56, + 4 days of July). Active = last_engaged >= 2026-07-27 AND is_former = false.
- Excluded from active counts: any is_former=true row; any row dated before 2026-07-27.
- Flag rules: active contacts < 2 (single-threaded), < 3 (under-threaded), or all active contacts in one persona.
- Only 2 rows failed the recency test: CT-A902AE (2026-06-01) and CT-913581 (2026-06-20).
- Result: 14 deals evaluated, 11 flagged, 3 pass (Deal-84DBA6: 3 active, 3 personas; Deal-4B0BEB: 4 active, 4 personas; Deal-D348E1: 5 active, 5 personas).

FLAGGED DEALS (11)

1. Deal-EC3025 (C-FDD0C7) — single-threaded
   Amount: not in data | Stage: not in data
   Active contacts: 1 (CT-047C54 champion 2026-09-02; CT-F2C1AE economic buyer EXCLUDED, is_former=true)
   Present: champion | Missing: economic buyer, HR admin, IT security, finance
   Add: economic buyer (not stage-ranked — stage missing)
   On-file unengaged fit: CT-6827DB (Chief People Officer, economic buyer)

2. Deal-92D97D (C-E23238) — single-threaded
   Amount: not in data | Stage: not in data
   Active contacts: 1 (CT-01F5B4 HR admin 2026-08-28; CT-A902AE champion STALE 2026-06-01)
   Present: HR admin | Missing: economic buyer, champion, IT security, finance
   Add: economic buyer
   On-file unengaged fit: none in unengaged_contacts.csv (stale champion CT-A902AE exists on the deal record itself, but no economic buyer on file)

3. Deal-50D386 (C-EB10E4) — under-threaded (2 < 3)
   Amount: not in data | Stage: not in data
   Active contacts: 2 (CT-AA41B2 champion 2026-09-01, CT-B9C35B HR admin 2026-08-25)
   Present: champion, HR admin | Missing: economic buyer, IT security, finance
   Add: economic buyer
   On-file unengaged fit: CT-A1C4B3 (Chief People Officer, economic buyer)

4. Deal-D0D6B5 (C-32918E) — under-threaded (all contacts one persona)
   Amount: not in data | Stage: not in data
   Active contacts: 3, all champions (CT-87CED4 2026-09-02, CT-DE6D7C 2026-08-19, CT-FD70B2 2026-08-07)
   Present: champion | Missing: economic buyer, HR admin, IT security, finance
   Add: economic buyer
   On-file unengaged fit: CT-1FA4DB (Chief People Officer, economic buyer)

5. Deal-5BFE3B (C-535D36) — under-threaded (2 < 3 AND all one persona)
   Amount: not in data | Stage: not in data
   Active contacts: 2, both champions (CT-57123B 2026-08-31, CT-5CE757 2026-08-12)
   Present: champion | Missing: economic buyer, HR admin, IT security, finance
   Add: economic buyer
   On-file unengaged fit: none on file

6. Deal-36C33F (C-077A0E) — single-threaded
   Amount: not in data | Stage: not in data
   Active contacts: 1 (CT-4FE556 IT security 2026-08-15; CT-405B45 champion and CT-86B22F economic buyer both EXCLUDED, is_former=true)
   Present: IT security | Missing: economic buyer, champion, HR admin, finance
   Add: economic buyer
   On-file unengaged fit: CT-1DB73E (Chief People Officer, economic buyer)

7. Deal-885F45 (C-5E8EFB) — under-threaded (2 < 3)
   Amount: not in data | Stage: not in data
   Active contacts: 2 (CT-51C81E economic buyer 2026-08-26, CT-D9A0E8 champion 2026-08-11)
   Present: economic buyer, champion | Missing: HR admin, IT security, finance
   Add: cannot rank by stage (stage missing); this is the only flagged deal with an active economic buyer
   On-file unengaged fit: CT-B3F25D (IT Security Lead, IT security)

8. Deal-FCBE5B (C-737030) — single-threaded
   Amount: not in data | Stage: not in data
   Active contacts: 1 (CT-4A5317 champion 2026-08-29)
   Present: champion | Missing: economic buyer, HR admin, IT security, finance
   Add: economic buyer
   On-file unengaged fit: none on file

9. Deal-5408B0 (C-2AE3AA) — under-threaded (2 < 3)
   Amount: not in data | Stage: not in data
   Active contacts: 2 (CT-D33AE4 champion 2026-09-01, CT-8742FD HR admin 2026-08-18)
   Present: champion, HR admin | Missing: economic buyer, IT security, finance
   Add: economic buyer
   On-file unengaged fit: CT-07FA76 (Chief People Officer, economic buyer)

10. Deal-C6D97A (C-5A8FC2) — under-threaded (all contacts one persona)
    Amount: not in data | Stage: not in data
    Active contacts: 3, all champions (CT-223DDC 2026-08-31, CT-B03555 2026-08-20, CT-4E8A2B 2026-08-05)
    Present: champion | Missing: economic buyer, HR admin, IT security, finance
    Add: economic buyer
    On-file unengaged fit: none on file

11. Deal-F9A08A (C-0D15DF) — single-threaded
    Amount: not in data | Stage: not in data
    Active contacts: 1 (CT-931B10 champion 2026-09-03; CT-913581 economic buyer STALE 2026-06-20)
    Present: champion | Missing: economic buyer, HR admin, IT security, finance
    Add: economic buyer
    On-file unengaged fit: CT-697541 (Chief People Officer, economic buyer); the stale EB CT-913581 on the deal record is also a re-engage candidate

Summary: 5 single-threaded (Deal-EC3025, Deal-92D97D, Deal-36C33F, Deal-FCBE5B, Deal-F9A08A), 4 under-threaded at 2 actives (Deal-50D386, Deal-5BFE3B, Deal-885F45, Deal-5408B0), 2 at 3 actives but a single persona (Deal-D0D6B5, Deal-C6D97A). Economic buyer missing on 10 of 11; no economic buyer on file anywhere for Deal-92D97D, Deal-5BFE3B, Deal-FCBE5B, Deal-C6D97A.
talk-track-from-transcriptsPASS 1.00
Plain-English question: From this rep's last 10 calls: how do they open, handle objections, and how often do they lock a next step?
Prompt — what the model saw
From this rep's last 10 calls: what they lead with in the first five minutes, how they handle the three most common objections, how often a concrete next step is agreed (rate), and every competitor a prospect raised. Quote at most one line per finding and only lines that appear in the transcripts. End with two coaching notes.
cost $0.0016 · 45s · in 63 / out 2,636 tok
✓ objection_count✓ next_step_rate✓ competitors✓ coaching✓ quote_used

✓ Correct Answer (ground truth)

Plain-English answer: 8 of 10 calls open with the same customer-win story; objection handling, next-step agreement rate, and every competitor raised, with quotes.
The answer in one line: n_transcripts = 10; objection_count = 3; rep = Alex Franklin (+3 more below)
n_transcripts10
objection_count3
allowed_new13 items — e.g. TT-001, TT-002, TT-003, TT-004
competitor_quotesAwardco=We're also in late talks with Awardco — their rewards catalog looks bigger than yours.; Kudos=How are you different from Kudos? Our CEO used them at her last company.
competitors_prospect_raisedAwardco, Kudos
next_stepagreed=7; total=10; rate=0.7
objectionsobjection=budget; count=4; quote=Honestly, budget is locked until next fiscal year — I can't add a new line item right now.; transcript_id=TT-001 | objection=timing; count=3; quote=This is good, but can we revisit it next quarter? Open enrollment starts in October and we're underwater.; transcript_id=TT-002 | objection=in-house; count=3; quote=We already do recognition with a spreadsheet and quarterly gift cards — why would we change?; transcript_id=TT-004
opener_themetheme=opens with a customer recognition win story (400-person retailer, 18% turnover reduction); count=8; total=10; quote=Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it.; transcript_id=TT-001
repAlex Franklin
rep_raised_competitors_excludedWorkhuman
graded checksobjection_count · next_step_rate · competitors · coaching · quote_used

Model's actual answer

show the model's full answer
**1. LEAD (first five minutes)**

8 of 10 calls open with the identical scripted retailer story — TT-001, TT-002, TT-003, TT-005, TT-006, TT-007, TT-008, TT-010 (8/10 = 80%):
> "Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards..."

The two exceptions tailor to the deal (TT-004, TT-009):
> "You asked for straight pricing last time, so let's start there." (TT-009, Deal-1E2498)

One rep-side competitor insert inside the first five minutes (TT-005, not prospect-raised):
> "And unlike Workhuman, our pricing includes the full rewards catalog with no extra margin."

**2. THREE MOST COMMON OBJECTIONS (prospect-raised) AND HANDLING**

Counting prospect objection lines across the 10 calls: budget-locked = 4 (TT-001, TT-003, TT-006, TT-010), revisit-next-quarter = 3 (TT-002, TT-005, TT-008), spreadsheet/status-quo = 3 (TT-004, TT-007, TT-009). Remaining stalls (committee x2, no-urgency x1, competitor-in-talks x2) are single/double occurrences.

- Budget locked (4x) → reframed as funded-from-savings with a $210k number:
> "Totally fair. Most teams fund this out of turnover savings — that retailer saved about $210k in avoided backfills, which is how their finance team signed off."
- Revisit next quarter / open enrollment (3x) → counters with a scoped pilot:
> "Makes sense. What if we scope a 90-day pilot with one department so you have internal data before next quarter's planning?"
- Spreadsheet status quo (3x) → pivots to automation + analytics:
> "Spreadsheets work until they scale — the difference is automation: milestones fire without HR lifting a finger, and you get analytics on who is being recognized."

Each handling line is delivered verbatim-identical every time it recurs.

**3. CONCRETE NEXT STEP AGREEMENT RATE**

Next step asked on 7 calls (the "working session... Thursday at 2pm" ask) and accepted on all 7: 7/7 = 100%. Never asked on TT-004 (Deal-403845), TT-007 (Deal-EDC141), TT-010 (Deal-84DBA6) — the rep conceded to committee/urgency stalls instead:
> "Understood — I'll leave it with you." (TT-004)

Agreement rate across all 10 calls: 7/10 = 70%.

**4. COMPETITORS RAISED BY PROSPECTS**

- Awardco — TT-003 (Deal-547B2B): "We're also in late talks with Awardco — their rewards catalog looks bigger than yours."
- Kudos — TT-007 (Deal-EDC141): "How are you different from Kudos? Our CEO used them at her last company."

Workhuman was raised only by the rep (TT-005), not the prospect. No other competitor appears in the transcripts.

**COACHING NOTES**

1. The ask works — use it on every call. When Alex proposes a concrete next step it is accepted 7 of 7 times; all three call-ends without an agreed step (Deal-403845, Deal-EDC141, Deal-84DBA6) are ones where he never made the ask after a committee or no-urgency stall. Script a response to "the committee decides" that lands a micro-committee step (e.g., a 20-min working session including the budget owner) rather than "I'll leave it with you."

2. The open and the objection handlers are being recited verbatim (8/10 identical opens; identical $210k and pilot lines on every recurrence), including the competitor deflects, which answer "bigger catalog" and "our CEO used Kudos" with the same automation-plus-analytics pivot and no proof. Tailor the case story to the prospect's industry/trigger and back each competitor response with one evidence point (comparison page, win rate, catalog/price detail) before late-stage deals against Awardco and Kudos harden.
reporting-analytics 5 tests
q3-weighted-forecast0.90
Plain-English question: Build the Q3 2026 weighted forecast from the open pipeline, and show the math deal by deal.
Prompt — what the model saw
Below are the open deals in the default sales pipeline with a close date between 2026-07-01 and 2026-10-15, pulled from the CRM warehouse on 2026-09-05. Stages: DS1 (discovery) through DS5 (contract). Build the Q3 2026 forecast (quarter = 2026-07-01 to 2026-09-30). Weighting rule: 100 percent of COMMIT plus 35 percent of BEST_CASE. PIPELINE deals count zero. Only deals with a close date inside the quarter count. Report: COMMIT total, BEST_CASE total, weighted forecast, count of deals in each category inside the quarter, the deals you excluded for being outside the quarter (count and total amount), and the top 5 BEST_CASE deals by amount inside the quarter. Show arithmetic. Then, in a final '## Data quality' section of no more than four sentences, name the data quality problems in this extract that would make the forecast unreliable if it were run unattended.
cost $0.0084 · 191s · in 5,977 / out 13,367 tok
✓ commit_total_44729✓ best_case_total_203565✓ weighted_115976.75✓ commit_count_7✓ best_case_count_24✓ excluded_count_32✗ excluded_amount_227575✓ top5_aliases✓ flags_blank_owner✓ flags_missing_why_buys

✓ Correct Answer (ground truth)

Plain-English answer: A stage-weighted Q3 forecast number computed from the open pipeline, with per-stage math that reconciles to the deal data.
The answer in one line: excluded_amount = 227575; best_case_total = 203565; weighted_forecast = 115,976.75; commit_total = 44729; excluded_out_of_quarter = 32; best_case_count = 24; flags = blank owner column + empty why_buys fields must be surfaced (+2 more below)
excluded_amount227,575
best_case_count24
best_case_total203,565
commit_count7
commit_total44,729
excluded_out_of_quarter32
weighted_forecast115,976.75
flagsblank owner column + empty why_buys fields must be surfaced
top5_best_caseDeal-2D7423, Deal-25F752, Deal-E53952, Deal-5EED42, Deal-FA32A0
graded checkscommit_total_44729 · best_case_total_203565 · weighted_115976.75 · commit_count_7 · best_case_count_24 · excluded_count_32 · excluded_amount_227575 · top5_aliases · flags_blank_owner · flags_missing_why_buys

Model's actual answer

show the model's full answer
Q3 2026 FORECAST (quarter window 2026-07-01 to 2026-09-30; rule: 100% COMMIT + 35% BEST_CASE, PIPELINE = 0)

COMMIT (in-quarter) — 7 deals, total 44,729
  Deal-547B2B 11,200 + Deal-B7EBD1 9,000 + Deal-403845 9,000 + Deal-A2B47C 6,360 + Deal-2465CE 5,400 + Deal-A5E80A 2,520 + Deal-499BF6 1,249 = 44,729

BEST_CASE (in-quarter) — 24 deals, total 203,565
  (sum of 24 deals closing 2026-09-10 through 2026-09-30)

WEIGHTED FORECAST
  44,729 + 0.35 × 203,565
  = 44,729 + 71,247.75
  = 115,976.75

Counts inside the quarter:
  COMMIT: 7 deals (44,729)
  BEST_CASE: 24 deals (203,565)
  PIPELINE: 41 deals — count zero by rule (not enumerated)

Excluded for being outside the quarter (COMMIT/BEST_CASE only): 10 deals, total 42,010
  Deal-D348E1 13,770 (COMMIT, 2026-10-15)
  Deal-C61CF7 5,400 (BEST_CASE, 2026-10-09)
  Deal-48B656 5,160 (2026-10-15)
  Deal-901332 3,600 (2026-10-15)
  Deal-47AE31 3,600 (2026-10-09)
  Deal-15D24F 3,600 (2026-10-09)
  Deal-ED725A 2,400 (2026-10-08)
  Deal-8AD4A5 1,800 (2026-10-07)
  Deal-5FDCE4 1,600 (2026-10-01)
  Deal-F5A622 1,080 (2026-10-08)
  (9 × 0.35 + 13,770 would have added ~16,463 to the weighted number if the window were misread)

Top 5 BEST_CASE deals in-quarter by amount:
  1. Deal-2D7423 — 38,935 (2026-09-30)
  2. Deal-25F752 — 24,000 (2026-09-25)
  3. Deal-E53952 — 19,656 (2026-09-30)
  4. Deal-5EED42 — 16,250 (2026-09-30)
  5. Deal-FA32A0 — 11,116 (2026-09-25)

## Data quality
Owner is blank on 85 of 86 rows, so an unattended run cannot attribute deals or detect single-rep optimism. Only 2 of the 31 COMMIT/BEST_CASE deals in the extract carry any why_buys narrative (33 show 0 chars), so category assignments are unverifiable against evidence. Deal-A5E80A is a DS1 and Deal-499BF6 a DS2 yet both are marked COMMIT, and the top five BEST_CASE deals (98,957 — 49% of the BC total) cluster on 2026-09-25/09-30, a classic quarter-end-boundary optimism pattern. The extract was pulled 2026-09-05 but still contains PIPELINE deals with close dates already in the past (e.g. Deal-333EBB 2026-08-28), confirming records are not being closed or aged out, so stage/category fields lag reality.
aha-moment-2x2PASS 1.00
Plain-English question: Do companies that hit 5+ givers and one redemption in month one actually retain better at 24 months?
Prompt — what the model saw
Test the activation hypothesis on the company cohort below: companies with 5+ unique givers (m1_users >= 5) AND 1+ successful redemption (m1_redemptions >= 1) in their first calendar month retain better at 24 months than those with one signal or neither. Every company in this extract is already 25+ months old, so current_status = 'active' means retained at 24 months. Report the full 2x2 (cohort size and 24-month retention rate for: both signals, givers-only, redemption-only, neither), state how many companies were excluded from the denominator and why (if none, say so), name the single signal with the largest retention lift, and state what this does and does not prove.
cost $0.0446 · 192s · in 26,031 / out 15,029 tok
✓ both_66✓ givers_47✓ redemp_31✓ neither_40✓ excluded_0✓ largest_lift_givers

✓ Correct Answer (ground truth)

Plain-English answer: The four cohort cells are graded against computed 24-month retention; the numbers come from the scorer's cohort table, not opinion.
The answer in one line: both_5plus_givers_and_redemption = 66.0% (31 of 47)
both_5plus_givers_and_redemption66.0% (31 of 47)
exclusionsnone — claim bounded as correlation, not causation
givers_only46.9% (23 of 49)
largest_single_signal_liftunique givers
neither40.0% (38 of 95)
redemption_only31.0% (9 of 29)
graded checksboth_66 · givers_47 · redemp_31 · neither_40 · excluded_0 · largest_lift_givers

Model's actual answer

show the model's full answer
Activation hypothesis test: retained = current_status exactly 'active'; cancelled + non_renewing = not retained. N = 220 rows, all classifiable, 0 excluded.

2x2 (24-month retention = active rate):

Both signals (m1_users >= 5 AND m1_redemptions >= 1): n = 47, active = 31. 31/47 = 65.96%
Givers-only (users >= 5, redemptions = 0): n = 49, active = 23. 23/49 = 46.94%
Redemption-only (users < 5, redemptions >= 1): n = 29, active = 9 (plus 2 non_renewing, 18 cancelled). 9/29 = 31.03%
Neither (users < 5, redemptions = 0): n = 95, active = 38 (plus 1 non_renewing, 56 cancelled). 38/95 = 40.00%

Check: 47 + 49 + 29 + 95 = 220. Actives: 31 + 23 + 9 + 38 = 101. Overall 101/220 = 45.91%.

Excluded from denominator: 0. No blank m1_users, m1_redemptions, or current_status; every row falls in exactly one cell.

Single signal with largest retention lift: givers (m1_users >= 5).
Marginal givers present: 54/96 = 56.25%; absent: 47/124 = 37.90%; lift = 18.35pp.
Marginal redemption present: 40/76 = 52.63%; absent: 61/144 = 42.36%; lift = 10.27pp.
Vs neither baseline (40.00%): givers-only +6.94pp (46.94 - 40.00); redemption-only -8.97pp (31.03 - 40.00).

Conclusion: directionally supports the hypothesis in this extract — both-signals retains highest at 66.0%, above givers-only, redemption-only, and neither.

What this does prove: descriptive association only, in this 220-company extract: first-month both-signals co-occur with higher 24-month active rate.

What this does not prove: causation. Observational with no randomization and no controls for confounders (product_tier, company size, industry, Slack/Teams/HRIS, signup month). Redemption-only n = 29 is small; 'active' is a current-status proxy per your definition, not a verified 24-month timestamp. Do not read as: causing 5+ givers plus 1+ redemption will raise retention.
arr-rollforward-reconciliationPASS 1.00
Plain-English question: Does CRM ARR match billing ARR — and if not, exactly where is the variance?
Prompt — what the model saw
Reconcile ARR between CRM company records and billing subscriptions as of 2026-09-05. Billing ARR is MRR x 12. Report both totals, the variance, and a decomposition into buckets (status mismatch, rounding, missing records, other) that sums exactly to the variance. List each mismatched account with a suggested owner. Business rule: any subscription with a term other than 12 months must have cf_agreement_end_date populated; list violations.
cost $0.0078 · 211s · in 2,058 / out 14,438 tok
✓ billing_total✓ crm_total✓ variance✓ buckets_sum✓ violations_listed

✓ Correct Answer (ground truth)

Plain-English answer: Billing $604,739.28 vs CRM $603,581.76: a $1,157.52 variance decomposed into status-mismatch, rounding, and missing-record buckets.
The answer in one line: n_companies = 40; n_mismatched_accounts = 7; as_of = 2026-09-05 (+4 more below)
n_companies40
n_mismatched_accounts7
as_of2026-09-05
billing_arr_total604739.28
bucket_sum_equals_variance1
buckets{"status_mismatch": {"amount": "-13158.48", "accounts": [{"company_alias": "C-0C8323BF", "billing_arr": "0.00", "hubspot_arr": "4905.24", "diff": "-4905.24", "suggested_owner": "Dana Mercer"}, {"compa
crm_arr_total603581.76
ruleterm_months != 12 requires cf_agreement_end_date
variance1157.52
variance_definitionbilling_arr_total - crm_arr_total (billing ARR = sum of MRR x 12 over status=active subscriptions)
violationssubscription_id=SUB-0002; company_alias=C-1794A52C; term_months=24 | subscription_id=SUB-0019; company_alias=C-22170CA1; term_months=36
graded checksbilling_total · crm_total · variance · buckets_sum · violations_listed

Model's actual answer

show the model's full answer
Reconciliation complete. All figures computed from the two files only; no owner/rep fields exist in either file, so suggested owners are by role (stated where relevant).

BASELINE (as of 2026-09-05)
- Billing ARR (MRR × 12, all 39 subs): $617,897.76
- Billing ARR, active subs only (37 subs): $604,739.28
  (= 617,897.76 − 408.77×12 − 687.77×12 = 617,897.76 − 4,905.24 − 8,253.24)
- CRM ARR (sum of 39 company records): $603,581.76
- Row counts: 39 subs, 39 CRM companies, but only 37 aliases overlap.

VARIANCE
CRM is a record of committed ARR, so I reconcile it against active-only billing:
  Variance = CRM − Billing(active) = 603,581.76 − 604,739.28 = **−$1,157.52**
(Reference: CRM − Billing(all 39) = 603,581.76 − 617,897.76 = −$14,316.00)

DECOMPOSITION (sums exactly to −$1,157.52)
1. Status mismatch: +$13,158.48
   SUB-000E / C-0C8323BF: cancelled in billing, still carries $4,905.24 in CRM (= 408.77×12, exact)
   SUB-000F / C-0DC4FB8C: cancelled in billing, still carries $8,253.24 in CRM (= 687.77×12, exact)
   CRM counts ARR for customers billing says are churned.
2. Missing records: −$11,952.00
   SUB-0004 / C-21629AA4: active in billing, ARR $28,449.24 (2,370.77×12), NO CRM company record → −28,449.24
   C-0D5BBE3A: in CRM at $16,497.24, NO billing subscription → +16,497.24
3. Rounding: +$36.00 (CRM values rounded to nearest $100)
   C-0D66DF9E: billed 1,932.00×12 = $23,184.00 vs CRM $23,200.00 → +16.00
   C-14D70CE0: billed 1,515.00×12 = $18,180.00 vs CRM $18,200.00 → +20.00
4. Other: −$2,400.00
   C-0F7269D7: billed 2,233.00×12 = $26,796.00 vs CRM $24,396.00 → −2,400.00
   (Implies CRM MRR of $2,033.00 vs billed $2,233.00 — a $200/mo delta; files don't say which side is current.)

Check: +13,158.48 − 11,952.00 + 36.00 − 2,400.00 = −$1,157.52 ✓
All other 32 overlapping aliases match billing MRR×12 to the cent.

MISMATCHED ACCOUNTS & SUGGESTED OWNER
(No ownership fields in the data; assignments below are by role.)
| Account | Issue | ARR impact | Suggested owner |
| C-0C8323BF (SUB-000E) | cancelled in billing, live ARR in CRM | +4,905.24 | RevOps data steward (close/zero CRM ARR) |
| C-0DC4FB8C (SUB-000F) | cancelled in billing, live ARR in CRM | +8,253.24 | RevOps data steward |
| C-21629AA4 (SUB-0004) | active sub, no CRM company | −28,449.24 | CRM admin + AE of record (create/attach record) |
| C-0D5BBE3A | CRM ARR, no subscription | +16,497.24 | Billing/Deal Desk (verify contract, or strip phantom ARR) |
| C-0D66DF9E | rounding | +16.00 | Record owner (sync CRM ARR to billing) |
| C-14D70CE0 | rounding | +20.00 | Record owner |
| C-0F7269D7 | $200/mo unexplained delta | −2,400.00 | RevOps + billing ops (rate-change audit) |

TERM-VIOLATION CHECK (term ≠ 12 requires cf_agreement_end_date)
Non-12-month subs: 4. Violations: 2
- SUB-0002 / C-1794A52C — 24 months, cf_agreement_end_date BLANK → VIOLATION
- SUB-0019 / C-22170CA1 — 36 months, cf_agreement_end_date BLANK → VIOLATION
- SUB-000C / C-0DB48281 — 24 months, populated (2027-11-30) — compliant
- SUB-001A / C-0FC4DBB8 — 36 months, populated (2027-11-30) — compliant
kvm-month-over-monthPASS 1.00
Plain-English question: How did the four core KVMs move in August versus July?
Prompt — what the model saw
Report the four core KVMs for 2026-08 against 2026-07: giving rate, redemptions per user, 1:1 meetings engagement, pulse check engagement. For each: value, prior value, absolute and relative change, direction. Name the KVM with the largest relative move and the segment (plan_tier or size_band) driving it if the data supports that; otherwise say the data does not. Under 200 words plus one table.
cost $0.0107 · 155s · in 12,386 / out 11,661 tok
✓ four_kvms✓ biggest_mover_named✓ direction_correct✓ segment_named✓ word_cap

✓ Correct Answer (ground truth)

Plain-English answer: Giving rate 60.27% (up 0.07%), redemptions per user 1.7302 (up 0.01%), with 1:1 meetings and pulse engagement each up fractionally — direction reported for all four.
The answer in one line: n_company_rows_per_month = 30; aggregation = simple mean of company-level values per month
n_company_rows_per_month30
aggregationsimple mean of company-level values per month
biggest_moverpulse_engagement
driving_segmentsize_band=enterprise; 2026-07=0.55; 2026-08=0.2743; rel_change_pct=-50.13
kvms{"giving_rate": {"2026-07": 0.6023, "2026-08": 0.6027, "abs_change": 0.0004, "rel_change_pct": 0.07, "direction": "up"}, "redemptions_per_user": {"2026-07": 1.73, "2026-08": 1.7302, "abs_change": 0.00
segment_detail{"smb": {"2026-07": 0.6588, "2026-08": 0.6573, "rel_change_pct": -0.23}, "mid_market": {"2026-07": 0.593, "2026-08": 0.5942, "rel_change_pct": 0.2}, "enterprise": {"2026-07": 0.55, "2026-08": 0.2743,
graded checksfour_kvms · biggest_mover_named · direction_correct · segment_named · word_cap

Model's actual answer

show the model's full answer
All 30 companies are tier_three (single plan tier), so the segment driver can only come from size_band. Values are unweighted means across companies (no headcount data provided to weight by).

Arithmetic (sum of 30 values ÷ 30 per month):
- Giving rate: Jul 0.6023 → Aug 0.6027 → +0.0004 abs; +0.0004/0.6023 = +0.07%
- Redemptions/user: Jul 1.7300 → Aug 1.7302 → +0.0002 abs; +0.01%
- 1:1 engagement: Jul 0.4469 → Aug 0.4472 → +0.0003 abs; +0.07%
- Pulse engagement: Jul 0.6006 → Aug 0.5086 → −0.0920 abs; −0.0920/0.6006 = −15.3%

| KVM | Jul-26 | Aug-26 | Abs Δ | Rel Δ | Dir |
|---|---|---|---|---|---|
| Giving rate | 0.6023 | 0.6027 | +0.0004 | +0.07% | up |
| Redemptions/user | 1.7300 | 1.7302 | +0.0002 | +0.01% | flat |
| 1:1 meetings engagement | 0.4469 | 0.4472 | +0.0003 | +0.07% | up |
| Pulse check engagement | 0.6006 | 0.5086 | −0.0920 | −15.3% | down |

Largest relative move: pulse engagement (−15.3%), driven by the enterprise size_band. Enterprise pulse fell 0.5500 → 0.2743 (−50.1%) — all 10 enterprise companies dropped 49–52% (e.g., C-0B2895EF 0.5619→0.2864, C-0D0B047C 0.5398→0.2619). Mid_market (+0.21%) and smb (−0.23%) barely moved; no non-enterprise company exceeded ±1.2%. The uniform ~50% drop across every enterprise account suggests a measurement/coverage change rather than organic behavior, but the data provided does not confirm a cause.
rewards-ytd-sectionPASS 1.00
Plain-English question: Produce the redemption section of the weekly rewards report through August.
Prompt — what the model saw
Produce the redemption section of the weekly rewards report through the last completed month (2026-08): redemption count, spend, unique redeemers, redemptions per redeemer, provider mix as percent of spend, top 5 countries by redemptions. Say which month is the last completed month. Provider shares must sum to 100.
cost $0.1455 · 65s · in 51,816 / out 17,741 tok
✓ count✓ spend✓ unique_redeemers✓ month_named✓ provider_mix✓ top_country

✓ Correct Answer (ground truth)

Plain-English answer: 378 redemptions, $27,846 spend, 235 unique redeemers at 1.61 each, custom rewards at 39% of spend, US leading all countries.
The answer in one line: redemption_count = 378; unique_redeemers = 235; redemptions_per_redeemer = 1.61; total_spend_usd = 27846.00
redemption_count378
redemptions_per_redeemer1.61
unique_redeemers235
provider_mix_pct_of_spendTangoCard=11.6; Tremendous=30.54; custom=39.05; Snappy=18.81
top5_countries_by_countcountry=US; count=244 | country=CA; count=24 | country=AU; count=21 | country=GB; count=17 | country=NL; count=17
total_spend_usd27846.00
graded checkscount · spend · unique_redeemers · month_named · provider_mix · top_country

Model's actual answer

show the model's full answer
Redemption section — YTD through last completed month: 2026-08

Scope: 2026-01 to 2026-08 inclusive. Last completed month is 2026-08; max redeemed_at in file is 2026-08-31, no later months present.

Redemption count: 378
45 + 36 + 50 + 52 + 43 + 46 + 58 + 48 = 378 for Jan through Aug

Spend: $27,846.00
10,873.00 + 8,505.00 + 5,238.00 + 3,230.00 = 27,846.00

Unique redeemers: 235 distinct user_key

Redemptions per redeemer: 1.61
378 / 235 = 1.6085

Provider mix, percent of spend:
custom: $10,873.00, 10,873 / 27,846 x 100 = 39.05%
Tremendous: $8,505.00, 8,505 / 27,846 x 100 = 30.54%
Snappy: $5,238.00, 5,238 / 27,846 x 100 = 18.81%
TangoCard: $3,230.00, 3,230 / 27,846 x 100 = 11.60%
Sum: 39.05 + 30.54 + 18.81 + 11.60 = 100.00

Top 5 countries by redemptions:
US: 244
CA: 24
AU: 21
GB: 17
NL: 17
Next: SG 12; DE, FR, CH 9 each. GB and NL tie at 17 for 4th/5th.
customer-success 5 tests
churn-save-eligibilityPASS 1.00
Plain-English question: Which at-risk accounts actually qualify for a churn-save offer, and which play fits each?
Prompt — what the model saw
Which at-risk accounts qualify for a churn-save offer under the documented eligibility rules, what amount is at stake per account and in total, and which play fits each (usage revival, executive touch, commercial concession)? Cite the signal that justifies each play. List accounts that look at risk but do not qualify and why.
cost $0.0080 · 139s · in 11,998 / out 11,622 tok
✓ eligible_set✓ total_at_stake✓ plays_cited✓ noneligible_named✓ rules_applied

✓ Correct Answer (ground truth)

Plain-English answer: Only accounts meeting all three rules qualify — e.g. C-0F6C0F34 at $49,707 (executive touch) and C-0B827671 at $25,365 (usage revival) — with amount and play per account.
The answer in one line: n_accounts = 30; snapshot_date = 2026-09-05
n_accounts30
eligibleaccount_alias=C-0F6C0F34; amount_at_stake=49707.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0B827671; amount_at_stake=25365.00; play=usage revival; justifying_signal=usage_trend_3m=declining over the last 3 months | account_alias=C-0B360C78; amount_at_stake=35748.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 75% - risk is commercial, not adoption | account_alias=C-0B0F1BAB; amount_at_stake=5494.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0CA21961; amount_at_stake=16829.00; play=usage revival; justifying_signal=seat utilization 26% is below 50% | account_alias=C-0E9C27D1; amount_at_stake=41235.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 85% - risk is commercial, not adoption … (+2 more)
non_eligible_at_riskaccount_alias=C-0BC71BDD; health_score=55 | account_alias=C-0BA71F12; health_score=52 | account_alias=C-0F6694C3; health_score=43 | account_alias=C-0BE96399; health_score=54 | account_alias=C-0F876796; health_score=47 | account_alias=C-0FCCD2DF; health_score=43 … (+1 more)
ruleshealth_score < 60, churn_save_eligible_amount > 0, renewal within 120 days of snapshot
snapshot_date2026-09-05
total_amount_at_stake224601.00
graded checkseligible_set · total_at_stake · plays_cited · noneligible_named · rules_applied

Model's actual answer

show the model's full answer
Eligibility screen (all three rules must pass; snapshot 2026-09-05, R3 window = renewal on or before 2027-01-03):

QUALIFYING ACCOUNTS — 8 total

| Account | Health | Churn-save $ at stake | ARR | Renewal (days) | Usage 3m | Seats used | Champion | Play |
|---|---|---|---|---|---|---|---|---|
| C-0F6C0F34 | 51 | $49,707.00 | $86,741 | 2026-10-03 (28d) | growing | 308/395 = 78.0% | no | Executive touch |
| C-0B827671 | 56 | $25,365.00 | $72,088 | 2026-11-14 (70d) | declining | 113/202 = 55.9% | yes | Usage revival |
| C-0B360C78 | 57 | $35,748.00 | $60,427 | 2026-10-28 (53d) | growing | 246/327 = 75.2% | yes | Commercial concession |
| C-0B0F1BAB | 38 | $5,494.00 | $15,391 | 2026-09-23 (18d) | flat | 238/363 = 65.6% | no | Executive touch |
| C-0CA21961 | 58 | $16,829.00 | $31,501 | 2026-12-28 (114d) | flat | 84/325 = 25.8% | yes | Usage revival |
| C-0E9C27D1 | 39 | $41,235.00 | $75,093 | 2026-09-24 (19d) | flat | 134/157 = 85.4% | yes | Commercial concession |
| C-0CEF69FD | 53 | $32,621.00 | $79,324 | 2026-11-21 (77d) | growing | 97/136 = 71.3% | no | Executive touch |
| C-0D3278C7 | 54 | $17,602.00 | $33,815 | 2026-11-12 (68d) | declining | 126/380 = 33.2% | yes | Usage revival |

Total at stake (churn-save eligible):
49,707 + 25,365 + 35,748 + 5,494 + 16,829 + 41,235 + 32,621 + 17,602 = $224,601.00
(Total ARR behind these 8: 86,741 + 72,088 + 60,427 + 15,391 + 31,501 + 75,093 + 79,324 + 33,815 = $454,380.00)

Play subtotals:
- Usage revival (3): 25,365 + 16,829 + 17,602 = $59,796.00
- Executive touch (3): 49,707 + 5,494 + 32,621 = $87,822.00
- Commercial concession (2): 35,748 + 41,235 = $76,983.00
- Check: 59,796 + 87,822 + 76,983 = $224,601.00 ✓

Play rationale (signal cited). Note: the provided files document eligibility rules only (R1–R3) — no documented play-assignment rules exist in the data. The mapping below is judgment applied to the provided columns, each tied to a cited signal:

- C-0F6C0F34 — executive touch: champion_active=false despite usage growing at 78.0% utilization. The risk signal is relationship loss, not usage; renewal in 28 days forces it now.
- C-0B827671 — usage revival: usage_trend_3m=declining and utilization down to 55.9% (113 of 202 seats). Champion exists to drive adoption internally.
- C-0B360C78 — commercial concession: usage growing at 75.2% utilization with an active champion — the only negative signal is health_score=57 and the $35,748 eligible amount, pointing to a value/commercial objection, not adoption.
- C-0B0F1BAB — executive touch: champion_active=false with the second-worst health in the file (38) and renewal in 18 days; usage is flat at 65.6%, so revival isn't the gap — sponsorship and urgency are.
- C-0CA21961 — usage revival: worst utilization of any qualifier, 25.8% (84 of 325 seats) with flat usage. Massive paid-but-unused seat base.
- C-0E9C27D1 — commercial concession: health 39 and renewal in 19 days, but usage is fully engaged (85.4% utilization, champion active, flat trend). Engaged-but-unhealthy at this proximity is a price/value conversation.
- C-0CEF69FD — executive touch: champion_active=false while usage grows at 71.3% — no usage problem to revive; the at-risk signal is the missing champion 77 days from renewal.
- C-0D3278C7 — usage revival: usage_trend_3m=declining and 33.2% utilization (126 of 380 seats); declining usage is the sole driver of its risk, and a champion is in place to run adoption.

AT-RISK BUT DO NOT QUALIFY — 7 accounts (health < 60, fails R2 and/or R3):

| Account | Health | Churn-save $ | Renewal (days from 2026-09-05) | Why excluded |
|---|---|---|---|---|
| C-0BC71BDD | 55 | $0.00 | 2026-10-27 (52d) | R2: eligible amount is 0 |
| C-0BA71F12 | 52 | $6,824.00 | 2027-04-11 (218d) | R3: renewal 98 days outside window |
| C-0F6694C3 | 43 | $0.00 | 2027-03-21 (197d) | R2 (amount 0) and R3 (outside window) |
| C-0BE96399 | 54 | $0.00 | 2026-10-29 (54d) | R2: eligible amount is 0 — despite declining usage at 27.9% utilization, no churn-save budget exists |
| C-0F876796 | 47 | $19,958.00 | 2027-02-06 (154d) | R3: renewal 34 days past the 2027-01-03 cutoff — otherwise the strongest save candidate ($19,958 eligible, declining, 23.2% utilization) |
| C-0FCCD2DF | 43 | $0.00 | 2027-04-23 (230d) | R2 and R3 both fail |
| C-10A56B0F | 54 | $0.00 | 2026-12-12 (98d) | R2: in window, but eligible amount is 0 |

Excluded-group context: combined ARR $276,707 (54,515 + 17,061 + 11,243 + 52,319 + 49,895 + 65,957 + 25,717), of which only $26,782 (6,824 + 19,958) is churn-save eligible but window-blocked. Two names — C-0F876796 and C-0BA71F12 — decay on every usage signal yet fail only on timing; worth flagging for an early re-screen when they enter the 120-day window.
expansion-kitPASS 1.00
Plain-English question: Build the expansion kit for the customer who just replied to our outreach.
Prompt — what the model saw
A customer replied to outreach (reply included). Build the expansion kit: seat coverage (licensed over headcount), two lines on usage health, headroom in seats and ARR at the current per-seat rate, who replied and whether they can buy, the right buyer from our contacts, and a reply email under 150 words that moves toward the right person without being pushy and cites one real usage fact.
cost $0.0015 · 36s · in 150 / out 2,375 tok
✓ coverage_pct✓ headroom_arr✓ right_buyer✓ replier_not_buyer✓ usage_fact✓ email_under_150

✓ Correct Answer (ground truth)

Plain-English answer: 37.5% seat coverage, 250 headroom seats worth $15,000 ARR at $60 a seat, usage up 43.2% in six months — and the replier cannot buy; Dana R. (VP People) is the right buyer.
The answer in one line: headroom_seats = 250; coverage_pct = 37.5; company_alias = C-0DDFC9A7 (+3 more below)
headroom_seats250
coverage_pct37.5
company_aliasC-0DDFC9A7
headroom_arr15000.00
per_seat_annual_rate60.00
replierMaria S.
replier_can_buy0
right_buyername=Dana R.; title=VP People
usage_factsactive_users_2026_08=126; active_users_2026_03=88; growth_pct_6m=43.2; licensed_seat_utilization_pct=84
usage_health_linesMonthly active users grew from 88 (2026-03) to 126 (2026-08), +43.2% over 6 months., Last month 126 of 150 licensed seats were active (84.0% utilization), approaching the seat cap.
graded checkscoverage_pct · headroom_arr · right_buyer · replier_not_buyer · usage_fact · email_under_150

Model's actual answer

show the model's full answer
EXPANSION KIT — C-0DDFC9A7

1) SEAT COVERAGE
- Licensed ÷ headcount: 150 ÷ 400 = 37.5% of employees are licensed
- Active usage vs. licenses: 126 ÷ 150 = 84% of paid seats active as of 2026-08

2) USAGE HEALTH (two lines)
- Monthly active users climbed every month: 88 (Mar) → 95 → 102 → 110 → 118 → 126 (Aug) = +38 users, 126/88 = +43.2% over 5 months (~7.4%/month compounded).
- At 84% of licensed seats active, they are within 24 seats of their cap — natural expansion pressure, not churn risk.

3) HEADROOM AT CURRENT PER-SEAT RATE
- Per-seat rate: $9,000 ARR ÷ 150 seats = $60/seat/year
- To license cap: 150 − 126 = 24 seats → 24 × $60 = $1,440 ARR
- To full headcount (400): 400 − 150 = 250 seats → 250 × $60 = $15,000 ARR incremental → $24,000 total ARR
- Trajectory: at ~7.4%/month (+9.4 users/mo avg), Aug's 126 users hit the 150-seat cap in ~2–3 months.

4) WHO REPLIED / BUYING AUTHORITY
- Maria S., People Operations Coordinator. Coordinator title = influencer/champion, not budget holder. No authority data provided — this is an inference from title, flagging it as such.

5) RIGHT BUYER
- Dana R., VP People — owns the people/engagement budget line, matches the product's buyer persona. (Sam K., Office Manager, last engaged 2025-11-03, is stale and not the economic buyer.)

6) REPLY EMAIL (90 words in body)

Subject: Re: your note — next step for C-0DDFC9A7

Hi Maria,

Thanks for replying — great to hear the team is getting real value out of Bonusly. The usage backs that up: active users grew from 88 in March to 126 in August, about 43% growth.

You're now using 126 of your 150 licensed seats, and we don't want adoption to hit a ceiling. Could you point me to the right person to approve an expansion — would that be Dana R.? Happy to make an intro, or I can reach out directly if you'd prefer.

No rush, and thanks again.

Best,
[Sender]

MISSING DATA (stated explicitly, not invented): contract renewal date, price tier history (the $60/seat is derived, ARR ÷ seats), confirmation of who holds budget at C-0DDFC9A7, and department-level usage.
onboarding-statusPASS 1.00
Plain-English question: What is done and what is missing in this customer's onboarding before the mid-onboarding call?
Prompt — what the model saw
Prep the CSM for the mid-onboarding call: what is complete (integrations connected, allowance set, admins added, first recognition given, first redemption), what is not, early engagement signals, and the three things to cover on the call. Mark nothing complete without a data field that shows it.
cost $0.0014 · 39s · in 229 / out 2,140 tok
✓ complete_items✓ incomplete_items✓ three_items✓ grounded

✓ Correct Answer (ground truth)

Plain-English answer: 25 days live: Slack connected, allowance set, admins added — but the HRIS integration is not connected, with early engagement signals and three call priorities.
The answer in one line: days_live = 25; company_alias = C-0D284E42
days_live25
call_itemsConnect the HRIS integration (integration_hris is empty) to automate user provisioning and offboarding., Get to first redemption (first_redemption_at is empty): walk admins through the rewards catalog and confirm allowance funding., Build the giving habit: daily active givers averaged 4.3 in week 1 vs 13.0 last week - agree on manager nudges and a recognition cadence.
checklist{"slack_integration_connected": {"complete": true, "field": "integration_slack", "value": "2026-08-12"}, "hris_integration_connected": {"complete": false, "field": "integration_hris", "value": ""}, "a
company_aliasC-0D284E42
engagement_signalsavg_active_givers_week1=4.3; avg_active_givers_last_week=13; latest_active_givers=15; days_with_activity=25
graded checkscomplete_items · incomplete_items · three_items · grounded

Model's actual answer

show the model's full answer
ONBOARDING STATUS — C-0D284E42 (signup 2026-08-11; usage data runs 8/11–8/18)

COMPLETE (field-backed)
- Slack integration: connected 2026-08-12 (day +1 after signup)
- Allowance set: 2026-08-13 (day +2)
- Admins: admins_added = 2. The field proves two exist; there is no target/required-admin field, so "admin setup complete" cannot be verified beyond the count.
- First recognition: sent 2026-08-15 14:22 (day +4)

NOT COMPLETE / NO EVIDENCE
- HRIS integration: integration_hris is blank — not done (data can't distinguish "not done" from "not recorded")
- First redemption: first_redemption_at is blank — no redemption shown, 3+ days after the first recognition

EARLY ENGAGEMENT SIGNALS
- active_givers by day (8/11→8/18): 3, 3, 4, 4, 5, 4, 7, 5
- Mean = (3+3+4+4+5+4+7+5)/8 = 35/8 = 4.4 givers/day
- Trend: opened at 3, closed the window at 5; peak 7 on 8/17 (+4 vs the 3 start = +133%). First two days avg 3.0 → last two days avg (7+5)/2 = 6.0 (doubled).
- Dip: 8/17→8/18 fell 7→5 (-2, -28.6% day-over-day).
- Recognition on 8/15 lines up with givers stepping 4→5, and the 7-peak lands 2 days after — early recognition appears to pull givers active.
- Caveat: only 8 days, no seat/user totals in the data, so giver penetration (% of org giving) cannot be computed.

THREE THINGS FOR THE CALL
1. Get the first redemption. Recognition fired 8/15 but first_redemption_at is still blank — walk through the redemption path live and confirm allowance balance/catalog access; first redemption is the missing activation event.
2. Resolve the HRIS blank. Decide connect-or-explicitly-defer (headcount sync, anniversary automations depend on it) rather than leaving it silently unchecked.
3. shore up giver momentum. 7 peaked then dropped to 5; agree one concrete push for the next two weeks (manager-led recognition or a kickoff campaign) and confirm what "enough admins" means — the file shows 2 with no target.

MISSING DATA (stated, not assumed)
- No post-8/18 usage rows despite today being 2026-09-25 — engagement trend beyond 8/18 is unknown.
- No required-admin count, no seat/user totals, no allowance balance field, no HRIS plan field.
renewal-risk-conflicting-datesPASS 1.00
Plain-English question: Two systems disagree on renewal dates — which do you trust per account, and who is actually renewing inside 90 days?
Prompt — what the model saw
Produce a 90-day renewal risk brief. Two systems hold renewal dates and disagree on some accounts; decide which to trust per account and say why (multi-year contracts are known to be wrong in ChurnZero). For every renewal: company, CSM, ARR, date used, seat utilization, 3-month usage trend, risk rating with one sentence of evidence. Flag every disagreement. Close with total ARR renewing and ARR at risk.
cost $0.0135 · 305s · in 5,407 / out 25,108 tok
✓ total_renewing✓ arr_at_risk✓ disagreements_flagged✓ trust_rule

✓ Correct Answer (ground truth)

Plain-English answer: Chargebee wins per account (multi-year contracts are known wrong in ChurnZero), with each trusted date, the reason, and a 90-day-window flag.
The answer in one line: n_accounts = 20; n_disagreements = 5; snapshot_date = 2026-09-05 (+2 more below)
n_accounts20
n_disagreements5
accounts20 items — e.g. account_alias=C-0B144C78; csm=Cole Ingram; arr=30899.00; trusted_renewal_date=2026-11-02; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.4; usage_3m_ratio=1.03; risk=low; evidence=3-month usage ratio 1.03 (last3 avg 103 vs prior3 100), seat utilization 75% | account_alias=C-0B20DB64; csm=Dana Mercer; arr=21770.00; trusted_renewal_date=2026-10-07; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=56.6; usage_3m_ratio=1; risk=medium; evidence=3-month usage ratio 1.00 (last3 avg 295 vs prior3 295), seat utilization 57% | account_alias=C-0B344485; csm=Elena Sinclair; arr=64384.00; trusted_renewal_date=2026-11-16; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=78; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 241 vs prior3 231), seat utilization 78% | account_alias=C-0B369871; csm=Dana Mercer; arr=85128.00; trusted_renewal_date=2026-10-29; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.1; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 330 vs prior3 316), seat utilization 75%
arr_at_risk_90d359409.00
disagreementsaccount_alias=C-0B7D2C30; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-15; is_multi_year=True; trusted=2026-09-15 | account_alias=C-0BBE3E60; cz_renewal_date=2027-09-26; cb_renewal_date=2026-09-26; is_multi_year=True; trusted=2026-09-26 | account_alias=C-0BCDB8C2; cz_renewal_date=2027-09-18; cb_renewal_date=2026-09-18; is_multi_year=True; trusted=2026-09-18 | account_alias=C-0D2AB865; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-22; is_multi_year=True; trusted=2026-09-22 | account_alias=C-0F5D2323; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-29; is_multi_year=True; trusted=2026-09-29
snapshot_date2026-09-05
total_arr_renewing_90d1048715.00
trust_rulemulti-year contracts: Chargebee is authoritative (ChurnZero known wrong); otherwise systems agree or Chargebee wins
window2026-09-05 to 2026-12-04
graded checkstotal_renewing · arr_at_risk · disagreements_flagged · trust_rule

Model's actual answer

show the model's full answer
90-DAY RENEWAL RISK BRIEF — as of 2026-09-25 (window: 2026-09-25 → 2026-12-24)

Note on identifiers: no company names exist in the source files; accounts are cited by their given aliases only.

1) WHICH SYSTEM TO TRUST

Rule applied: where ChurnZero (CZ) and Chargebee (CB) disagree, the 5 disagreements are ALL on accounts with term_months > 12 (is_multi_year=true). Per the known CB defect on multi-year contracts, ChurnZero is trusted for those. For the 15 twelve-month accounts the two systems agree exactly, so no adjudication was needed.

Evidence for trusting CZ on the multi-year accounts:
- C-0BCDB8C2 (36-mo): CZ 2027-09-18 = CB 2026-09-18 + exactly 12 months — CB appears frozen at a stale pre-extension/anniversary date.
- C-0BBE3E60 (24-mo): CZ 2027-09-26 = CB 2026-09-26 + exactly 12 months — same signature.
- C-0B7D2C30, C-0D2AB865, C-0F5D2323: CB dates (2026-09-15, 09-22, 09-29) are not reproducible from CZ dates by any whole-term arithmetic; CZ trusted as the CS system of record, CB flagged stale.

⚠ DISAGREEMENT FLAGS (all 5, CZ used):
| Alias | CZ date (USED) | CB date (rejected) | Term |
| C-0B7D2C30 | 2026-09-10 | 2026-09-15 | 36mo |
| C-0BCDB8C2 | 2027-09-18 | 2026-09-18 | 36mo |
| C-0D2AB865 | 2026-09-10 | 2026-09-22 | 24mo |
| C-0BBE3E60 | 2027-09-26 | 2026-09-26 | 24mo |
| C-0F5D2323 | 2026-09-10 | 2026-09-29 | 24mo |
All other 15 accounts: CZ = CB, 12-month terms.

2) RENEWALS (utilization = seats_used/seats; trend = Jun–Aug 2026 avg vs Mar–May 2026 avg active users)

HIGH RISK — $273,989 ARR
- C-0B7D2C30 | Dana Mercer | $65,901 | 2026-09-10 (CZ, ⚠flag; date already PASSED) | util 57.6% (274/476) | trend −18.2% ((97+94+84)/3=91.7 vs (119+110+107)/3=112.0) — steep 12-month usage slide into a renewal date that has already lapsed.
- C-0D2AB865 | Elena Sinclair | $38,022 | 2026-09-10 (CZ, ⚠flag; PASSED) | util 61.4% (250/407) | trend −18.9% (116.7 vs 144.3) — same lapsed-date + collapsing usage combination.
- C-0F5D2323 | Cole Ingram | $90,647 | 2026-09-10 (CZ, ⚠flag; PASSED) | util 28.5% (111/390) | trend +3.5% (19.7 vs 19.0) — flat usage of only ~19 people against 390 seats makes the $90.6K contract grossly over-scoped at renewal.
- C-0EC6999D | Elena Sinclair | $79,419 | 2026-10-03 (both agree; 8 days out) | util 27.7% (31/112) | trend +6.7% (16.0 vs 15.0) — renewing within two weeks at under 30% seat utilization.

MEDIUM RISK — $275,349 ARR
- C-0BCDB8C2 | Cole Ingram | $54,427 | 2027-09-18 (CZ, ⚠flag; OUTSIDE 90-day window) | util 54.7% (232/424) | trend −17.6% (118.3 vs 143.7) — heavy decline and low utilization, but the corrected CZ date defers the renewal a full year.
- C-0BBE3E60 | Dana Mercer | $30,993 | 2027-09-26 (CZ, ⚠flag; OUTSIDE window) | util 64.9% (74/114) | trend −19.5% (39.0 vs 44.3) — worst decline in the book, yet 12 months out per the trusted CZ date.
- C-0B20DB64 | Dana Mercer | $21,770 | 2026-10-07 | util 56.6% (214/378) | trend +0.1% (295.3 vs 295.0) — usage pinned to ~295 active vs 378 seats: persistent ~120-seat excess.
- C-0BBC4E7A | Cole Ingram | $56,374 | 2026-10-10 | util 67.7% (228/337) | trend −0.9% (140.7 vs 142.0) — flat but mid-tier utilization days before renewal.
- C-0FD551AB | Elena Sinclair | $48,815 | 2026-10-14 | util 55.9% (210/376) | trend −1.6% (123.7 vs 125.7) — 166 unused seats, no growth to absorb them.
- C-0F9F8F13 | Dana Mercer | $46,230 | 2026-10-18 | util 56.5% (199/352) | trend +0.2% (184.0 vs 183.7) — static usage at barely over half of licensed seats.
- C-0BC34584 | Cole Ingram | $16,740 | 2026-10-22 | util 66.2% (327/494) | trend +1.0% (104.7 vs 103.7) — 167 idle seats against flat ~104-user activity.

LOW RISK — $499,377 ARR
- C-0B7A7546 | Elena Sinclair | $35,062 | 2026-10-25 | util 88.8% (182/205) | trend +4.3% (64.0 vs 61.3) — near-full adoption and rising.
- C-0B369871 | Dana Mercer | $85,128 | 2026-10-29 | util 75.1% (317/422) | trend +4.3% (329.7 vs 316.0) — steady 12-month growth (289→333).
- C-0B144C78 | Cole Ingram | $30,899 | 2026-11-02 | util 75.4% (169/224) | trend +3.0% (102.7 vs 99.7) — growing usage on healthy utilization.
- C-0FC4DBB8 | Elena Sinclair | $94,732 | 2026-11-05 | util 76.7% (356/464) | trend +4.2% (191.0 vs 183.3) — largest contract in file, consistently expanding.
- C-0D5BBE3A | Dana Mercer | $39,740 | 2026-11-09 | util 83.3% (85/102) | trend +3.9% (89.7 vs 86.3) — high and rising engagement.
- C-0FB9D5AF | Cole Ingram | $63,158 | 2026-11-13 | util 72.4% (144/199) | trend +4.6% (174.0 vs 166.3) — strongest growth rate in the book.
- C-0B344485 | Elena Sinclair | $64,384 | 2026-11-16 | util 78.0% (224/287) | trend +4.2% (240.7 vs 231.0) — monotonic 12-month growth (211→244).
- C-0CB2C1B4 | Dana Mercer | $40,628 | 2026-11-20 | util 81.6% (386/473) | trend −2.0% (48.0 vs 49.0) — high utilization; the −2% wobble is 1 user/month of noise.
- C-22170CA1 | Cole Ingram | $45,646 | 2026-11-24 | util 85.4% (251/294) | trend +2.6% (145.7 vs 142.0) — strong adoption trajectory.

3) TOTALS (arithmetic)

File total ARR: $1,048,715 (sum of all 20 accounts).
- Due in the 90-day window (dates 2026-10-03 → 2026-11-24, both systems agreeing): 79,419+21,770+56,374+48,815+46,230+16,740+35,062+85,128+30,899+94,732+39,740+63,158+64,384+40,628+45,646 = $768,725 renewing.
- Already lapsed per trusted CZ dates (2026-09-10 ×3): 65,901+38,022+90,647 = $194,570 — status (renewed/expanded/churned) is NOT in the provided data; needs confirmation.
- Outside the window (CZ-corrected 2027 dates): 54,427+30,993 = $85,420.

ARR at risk:
- HIGH: $273,989 = 65,901+38,022+90,647+79,419 — of which only $79,419 (C-0EC6999D) is genuinely ahead of us in the window; $194,570 is already past date.
- MEDIUM: $275,349 (high+medium combined = $549,338 = 52.4% of total book, 1048715 → 549338/1048715).
- HIGH+MEDIUM actually inside the 90-day window: 269,348 = 79,419 + 189,929 (medium minus the two 2027 accounts: 275,349−85,420=189,929).
- LOW: $499,377. Check: 273,989+275,349+499,377 = $1,048,715 ✓

Data gaps stated explicitly: no renewal outcomes for the three lapsed 2026-09-10 dates; no company names (aliases only); no NPS/health-score or contract-value-per-seat data to refine the ratings beyond utilization + trend.
ticket-theme-synthesis0.80
Plain-English question: What are the real themes in this quarter's support tickets, ignoring the existing tags?
Prompt — what the model saw
Synthesize the quarter's support tickets into 4 to 7 themes. Do not trust existing tags; read the text. Per theme: name, count, share, distinct accounts, ARR affected, two ticket ids, one-line recommendation. Separate broad patterns from single-account noise and rank by ARR exposure, not volume.
cost $0.0070 · 132s · in 12,457 / out 9,441 tok
✗ theme_count_range✓ counts_verifiable✓ single_account_flagged✓ arr_ranked✓ ticket_ids

✓ Correct Answer (ground truth)

Plain-English answer: 80 tickets in 4-7 themes — points-not-posting leads (20 tickets, 25%, $31,100 ARR affected) — each with counts, distinct accounts, ARR, and a recommendation.
The answer in one line: total_tickets = 80; quarter_window = 2026-06-01 through 2026-08-31
total_tickets80
noteexisting tags are noisy/unreliable; classify by text
quarter_window2026-06-01 through 2026-08-31
ranking_by_arr_exposurehris_provisioning, redemption_failures, billing_dispute, points_not_posting, slack_sync
sample_ticket_ids80 items — e.g. IC-460004, IC-460016, IC-460001, IC-460041
single_account_noise_themebilling_dispute
themes{"points_not_posting": {"count": 20, "share_pct": 25.0, "distinct_accounts": 9, "arr_affected": "31100.00", "example_ticket_ids": ["IC-460001", "IC-460002"], "single_account": false}, "slack_sync": {"
graded checkstheme_count_range · counts_verifiable · single_account_flagged · arr_ranked · ticket_ids

Model's actual answer

show the model's full answer
THEME SYNTHESIS — Intercom support tickets, 2026-06-01 → 2026-08-29
Base: 80 tickets, 24 distinct accounts, total account ARR $284,800.
Method: classified on body_text only. Tags are demonstrably unreliable (e.g., all 16 invoice-error tickets carry tags of feedback/question/urgent/bug/how-to, none consistent; HRIS tickets tagged "billing"; points tickets tagged "billing"). Every account maps to exactly one theme, so theme ARR sums add to the $284,800 total with no double counting: 3+7+1+9+4 = 24 accounts; 114,000+68,800+52,000+31,100+18,900 = 284,800 ✓; 12+18+16+20+14 = 80 tickets ✓.

RANKED BY ARR EXPOSURE (not volume):

1) HRIS PROVISIONING SILENTLY FAILING — BROAD, HIGH-ARR
   Count 12 (12/80 = 15.0%) | Accounts 3: C-0B2213A9 ($36,000), C-0F6C0F34 ($30,000), C-0DDFC9A7 ($48,000)
   ARR affected: 36,000 + 30,000 + 48,000 = $114,000 (40.0% of ticketed ARR; 114,000/284,800)
   Example ids: IC-460062, IC-460060
   Recommendation: treat as sev-1 — sync "skips" hires with zero errors in the provisioning log, so failures are invisible; build a failed/skipped-provisioning alert and audit the three accounts' unprovisioned new hires this week.

2) GIFT-CARD REDEMPTION FAILURES + POINTS DEDUCTED WITHOUT DELIVERY — BROAD
   Count 18 (18/80 = 22.5%) | Accounts 7: C-0B0F1BAB, C-0B827671, C-0CEF69FD, C-0D9CA315, C-0F876796, C-0FCCD2DF, C-14264ABD (ARR 10,300+10,700+8,900+9,600+8,700+9,600+11,000)
   ARR affected: $68,800 (24.2%)
   Example ids: IC-460024, IC-460032
   Recommendation: the checkout hang plus "order errored but points still deducted" is a ledger-integrity bug — make redemptions idempotent with automatic point refund on fulfillment failure, and recredit the affected orders.

3) INVOICE / RENEWAL BILLING ERRORS — SINGLE-ACCOUNT NOISE, NOT A PRODUCT THEME
   Count 16 (16/80 = 20.0%) | Accounts 1: C-0E9C27D1
   ARR affected: $52,000 (18.3%)
   Example ids: IC-460071, IC-460078
   Recommendation: everything is the same account repeating an uncorrected seat-count error ("third invoice in a row") and wrong renewal tier — assign one owner to fix the billing record and issue corrected invoices; this is a churn-risk escalation on your largest single account, not a systemic billing defect.

4) RECOGNITION POINTS NOT POSTING — BROAD BUT LOW-ARR
   Count 20 (20/80 = 25.0%) — highest volume | Accounts 9: C-0B2895EF, C-0BE96399, C-0BF20542, C-0D0B047C, C-0D284E42, C-0D3278C7, C-0D6CC8E3, C-0DD0626C, C-21FEBCBB
   ARR affected: $31,100 (10.9%)
   Example ids: IC-460004, IC-460001
   Recommendation: recognitions report "delivered" but balances never credit, often fleet-wide after weekends — reconcile the delivery-to-credit queue and add a consumer-lag monitor; broad trust damage to the core loop despite small accounts.

5) SLACK INTEGRATION BROKEN — BROAD, LOW-ARR
   Count 14 (14/80 = 17.5%) | Accounts 4: C-0B843542, C-0BA71F12, C-10A56B0F, C-8C2E8F00 (ARR 4,400+3,900+5,400+5,200)
   ARR affected: $18,900 (6.6%)
   Example ids: IC-460049, IC-460047
   Recommendation: re-auth "does not stick" and sync toggles reset themselves — fix Slack token/session persistence server-side so reconnections survive, rather than asking users to re-authorize repeatedly.

BROAD PATTERNS vs SINGLE-ACCOUNT NOISE:
- Truly broad: themes 1, 2, 4, 5 (3–9 accounts each). Themes 4 and 5 are volume-heavy but carry the least revenue; themes 1 and 2 combine breadth with real ARR and should lead the fix order.
- Single-account noise: theme 3 is 20% of all ticket volume from one account (C-0E9C27D1 filed 16 tickets — 20% of the quarter's total) inflating the count side of any naive ranking; it matters only as a $52,000 retention risk.
- Data gap: no resolution/closure status or per-invoice dollar amounts are in the file, so I cannot quantify actual revenue lost (e.g., the 200-vs-150-seat overcharge) — only contract ARR at risk.
marketing 5 tests
case-study-match0.75
Plain-English question: Which three existing customers are the best case-study matches for this prospect?
Prompt — what the model saw
Given the prospect, rank the three most similar existing customers with public case studies for use as social proof. Similarity on industry, size band, use case, region. Cite the fields that drove each match. Do not name any customer without a case study and never include billing data or contact names.
cost $0.0011 · 26s · in 930 / out 1,402 tok
✗ top3✓ decoys_excluded✓ cites_fields✓ no_billing

✓ Correct Answer (ground truth)

Plain-English answer: Three ranked matches on industry, size, use case, and region (e.g. C-11C31562), with the higher-scoring decoy explicitly excluded for a mismatched field.
The answer in one line: prospect_alias = C-82AF3719
constraintsno_billing_data, no_contact_names, only_customers_with_has_case_study_true_named
decoys_excludedcompany_alias=C-D6217CAA; similarity_score=4; reason=no_case_study | company_alias=C-C153868F; similarity_score=4; reason=no_case_study
prospect_aliasC-82AF3719
prospect_fields_usedindustry, size_band, use_case, region
top3company_alias=C-11C31562; similarity_score=3 | company_alias=C-64171065; similarity_score=3 | company_alias=C-A13C193D; similarity_score=3
graded checkstop3 · decoys_excluded · cites_fields · no_billing

Model's actual answer

show the model's full answer
Blocked: the prospect record is missing. No prospect fields (industry, size_band, use_case, region) were provided in the input — only customers.csv. Similarity cannot be scored or ranked without them, and I won't guess a prospect profile.

What the data does support — the eligible social-proof pool (has_case_study = true), 7 of 20 rows:

| alias | industry | size_band | use_case | region |
|---|---|---|---|---|
| C-FEFCC020 | Education | Enterprise | wellness | NA-East |
| C-11209339 | Manufacturing | SMB | peer_milestones | APAC |
| C-9D69D147 | Financial Services | Mid-Market | retention | NA-West |
| C-11C31562 | Manufacturing | Mid-Market | employee_recognition | NA-West |
| C-CD4829A7 | Technology | Mid-Market | retention | NA-East |
| C-0CE315B2 | Education | SMB | wellness | EMEA |
| C-64171065 | Technology | Mid-Market | employee_recognition | NA-East |

Arithmetic: 20 rows total; 13 have has_case_study=false and are excluded; 7 remain eligible.

To produce the top-3 ranking, provide the prospect's four fields. I'd then score each eligible customer as exact matches per field (industry, size_band, use_case, region; max 4), rank descending, and cite the matching fields per pick.
channel-attribution-mathPASS 1.00
Plain-English question: Compute cost per SQM, cost per SQO, and pipeline per dollar for every channel over six months.
Prompt — what the model saw
Compute channel performance for the trailing 6 months. Paid: spend, SQMs, SQOs, cost per SQM, cost per SQO, SQM-to-SQO rate, pipeline amount, pipeline per dollar. Organic: volume, SQO rate, pipeline. Report a channel with spend and zero SQMs as undefined, not zero. Flag rows where SQO date precedes SQM date. Finish with a reallocation recommendation and your confidence given sample sizes.
cost $0.0043 · 120s · in 133 / out 8,383 tok
✓ undefined_channel✓ undefined_not_zero✓ flagged_rows✓ best_channel✓ reallocation

✓ Correct Answer (ground truth)

Plain-English answer: LinkedIn Ads at $960/SQM and $3,000/SQO, organic search at $90,000 pipeline on zero spend, paid social flagged at $18,000 with zero SQMs, and a channel winner named.
The answer in one line: best_pipeline_per_dollar_channel = paid_search
best_pipeline_per_dollar_channelpaid_search
confidence_notewebinars and referral have small SQM samples (<20); treat rates as directional.
flagged_rows_sqo_before_sqmCT-000041, CT-000044
months2026-03, 2026-04, 2026-05, 2026-06, 2026-07, 2026-08
per_channel{"linkedin_ads": {"spend_usd": 24000, "sqms": 25, "sqos": 8, "pipeline_amount": 96000, "cost_per_sqm": 960.0, "cost_per_sqo": 3000.0, "sqm_to_sqo_rate": 0.32, "pipeline_per_dollar": 4.0}, "organic_sea
reallocation_recommendationfrom=paid_social; to=paid_search; rationale=paid_social has spend but zero SQMs (cost per SQM undefined, not zero); paid_search has the highest pipeline per dollar.
undefined_channelspaid_social
graded checksundefined_channel · undefined_not_zero · flagged_rows · best_channel · reallocation

Model's actual answer

show the model's full answer
CHANNEL PERFORMANCE — TRAILING 6 MONTHS (2026-03 → 2026-08)

Scope note: all 122 contact rows fall inside the window (dates 2026-03-01 to 2026-08-28); spend file covers the same 6 months. Organic channels (organic_search, referral) have no rows in channel_spend.csv — treated as zero-spend. paid_social has spend but zero contact rows.

=== PAID CHANNELS ===

paid_search
  Spend:      6000 × 6 = $36,000
  SQMs:       40
  SQOs:       18
  Cost/SQM:   36,000 / 40 = $900
  Cost/SQO:   36,000 / 18 = $2,000
  SQM→SQO:    18 / 40 = 45.0%
  Pipeline:   18 × 40,000 = $720,000
  Pipe/$:     720,000 / 36,000 = 20.0

linkedin_ads
  Spend:      4000 × 6 = $24,000
  SQMs:       25
  SQOs:       8
  Cost/SQM:   24,000 / 25 = $960
  Cost/SQO:   24,000 / 8 = $3,000
  SQM→SQO:    8 / 25 = 32.0%
  Pipeline:   8 × 12,000 = $96,000
  Pipe/$:     96,000 / 24,000 = 4.0

paid_social
  Spend:      3000 × 6 = $18,000
  SQMs:       0  → cost/SQM = UNDEFINED (0 SQMs, not $0)
  SQOs:       0  → cost/SQO = UNDEFINED; SQM→SQO rate = UNDEFINED
  Pipeline:   $0; Pipe/$ = 0 / 18,000 = 0.0

webinars
  Spend:      1500 × 6 = $9,000
  SQMs:       12
  SQOs:       5
  Cost/SQM:   9,000 / 12 = $750
  Cost/SQO:   9,000 / 5 = $1,800
  SQM→SQO:    5 / 12 = 41.7%
  Pipeline:   5 × 12,000 = $60,000
  Pipe/$:     60,000 / 9,000 = 6.7

PAID TOTAL: spend $87,000 | SQMs 77 | SQOs 31 (31/77 = 40.3%) | pipeline $876,000
  blended cost/SQM = 87,000/77 = $1,130 | cost/SQO = 87,000/31 = $2,806 | Pipe/$ = 876,000/87,000 = 10.1

=== ORGANIC CHANNELS (no spend data available) ===

organic_search
  Volume:     30 SQMs | SQOs 10 → SQO rate 10/30 = 33.3%
  Pipeline:   10 × 9,000 = $90,000

referral
  Volume:     15 SQMs | SQOs 6 → SQO rate 6/15 = 40.0%
  Pipeline:   6 × 8,000 = $48,000

ORGANIC COMBINED: 45 SQMs, 16 SQOs (35.6%), pipeline $138,000

=== DATA QUALITY FLAGS ===

SQO date precedes SQM date (both linkedin_ads, $12,000 pipeline each):
  - CT-000044: SQM 2026-07-23, SQO 2026-07-18 (-5 days)
  - CT-000041: SQM 2026-06-14, SQO 2026-06-09 (-5 days)
  Sensitivity if excluded: linkedin_ads SQOs 6, rate 6/25 = 24.0%, cost/SQO 24,000/6 = $4,000, pipeline $72,000, Pipe/$ = 3.0.
Also noted: CT-000007 has same-day SQM/SQO (2026-03-28) — plausible, not flagged. paid_social's zero rows may be a tracking gap rather than true zero performance; the data cannot distinguish.

=== RECOMMENDATION ===

Shift spend out of linkedin_ads and paid_social toward paid_search. paid_search converts 20x faster per dollar than linkedin_ads (20.0 vs 4.0 Pipe/$) and at a $2,000 cost/SQO vs $3,000 — and that's before the two flagged linkedin rows, whose exclusion would worsen linkedin to ~3.0 Pipe/$. paid_social produced zero attributable leads against $18,000 spend; pause it pending an attribution audit before concluding it's dead. webinars is the efficiency story on rate (41.7%, $1,800 cost/SQO — best in paid) but tiny volume; worth a small test increase rather than a reallocation target. Caveat: paid_search's Pipe/$ edge is partly deal size ($40k avg vs $12k elsewhere), so the gap is attribution + ASP, not conversion alone.

Confidence: MODERATE on the headline call (paid_search > linkedin_ads; n=40 vs n=25 SQMs, consistent 6-month spend). LOW on webinars scaling (n=12 SQMs, 5 SQOs) and on paid_social being truly zero rather than mis-tracked (n=0). First-touch-only attribution means cross-channel assists are invisible — treat magnitudes, not ratios, as directional.
competitor-battlecard-updatePASS 1.00
Plain-English question: Update the Rivally battlecard with sourced pricing and where each side wins.
Prompt — what the model saw
Update the battlecard for the competitor. Sections: one-line positioning, pricing with source and date (newer source wins, note the conflict), where they win, where we win, objections and responses, recent changes, our 12-month win/loss record against them. Cite a snippet id for every factual claim. Rep opinion on a call is not a fact about the competitor. Anything from the old card you cannot re-source gets marked unverified.
cost $0.0042 · 129s · in 248 / out 8,130 tok
✓ competitor_named✓ pricing✓ win_loss✓ unverified_marked✓ citations

✓ Correct Answer (ground truth)

Plain-English answer: Rivally at $7 per user per month (the August source wins over the stale $5, conflict noted), a 13-7 win/loss record, and two unverified old claims flagged.
The answer in one line: competitor = Rivally
citation_ids_available25 items — e.g. S01, S02, S03, S04
competitorRivally
pricingvalue_per_user_per_month=7; source_snippet=S17; source_date=2026-08-12
rep_opinion_snippet_ids_not_factsS09, S21
unverified_old_claimsRivally lacks a Slack integration., Rivally was acquired by WorkHuman in 2025.
win_loss_12mowins=13; losses=7
graded checkscompetitor_named · pricing · win_loss · unverified_marked · citations

Model's actual answer

show the model's full answer
# BATTLECARD: RIVALLY — updated 2026-09-25

## 1. Positioning (one line)
Points-based recognition for mid-market (S02, S04), pushing upmarket into EU enterprise (S11, S12, S15) with a lightweight engagement-survey add-on, "Rivally Pulse" (S06, S23).

## 2. Pricing
- Current list: Recognition Starter $7 per user/month, annual billing required — S17, pricing_page, 2026-08-12.
- Conflict note: earlier pricing_page snapshots showed $5/user/mo (S03, 2026-01-20; S08, 2026-04-01). Newer source wins: $7 is current; $5 was accurate as of Jan–Apr 2026. Increase = $7 − $5 = +$2/user/mo = +40% (2/5).
- Observed deal quotes (call_notes deal mentions — secondhand, not list price):
  - $6.50/user/mo to a 500-seat prospect, annual term — S13, 2026-06-02. Implied ACV: 500 × 6.50 × 12 = $39,000/yr (list at $7 would be 500 × 7 × 12 = $42,000/yr).
  - $7 list less 15% for a 3-year term — S18, 2026-08-14. Effective: 7 × 0.85 = $5.95/user/mo.
- Excluded: AE opinion that Rivally is "discounting aggressively" (S21) is rep opinion, not a pricing fact.
- Missing: Bonusly's own pricing is not in the provided data — no price-differential claim can be made here.

## 3. Where they win
- Recognition feed engagement — praised by reviewers (S02, 2025-12-15; S16, 2026-07-19).
- Fast time-to-value: setup under a week, Slack integration worked out of the box (S04, 2026-02-02). This corrects the old card (see §8).
- EU: multi-language support and distributed-EU strength praised by an EU enterprise reviewer (S12, 2026-05-21); EU data residency GA + Dublin office (S15, 2026-07-01); ex-Workday VP EMEA hired to lead expansion (S11, 2026-05-09); pitched EU data residency in an active competitive eval (S05, 2026-02-18, deal mention).
- Support: response time under 4 hours praised (S22, 2026-08-30).

## 4. Where we win
- Analytics depth: their analytics "limited" (S02), dashboards "basic compared to enterprise tools" (S07, 2026-03-22); an 800-seat prospect chose Bonusly over Rivally citing analytics depth (S25, 2026-09-03, deal mention).
- Enterprise admin: no SCIM provisioning, manual user management painful (S10, 2026-04-28); admin tooling "lags peers" (S16); no bulk recognition editing (S24, 2026-09-02).
- EMEA caveat to their EU story: rewards catalog in EMEA thinner than US (S14, 2026-06-14).
- Displacement angle: off-Rivally migration was hard due to CSV-only analytics exports (S20, 2026-08-25) — their lock-in, and their data-loss risk on switch.

## 5. Objections and responses
- "Rivally is cheaper." Response: their list moved $5 → $7 (S03 → S17); discounting seen at $6.50 quoted (S13) and 15% off 3-year = $5.95 effective (S18). Do not overstate: S21 (AE "discounting aggressively") is opinion, not usable. Gap: our price point isn't in the provided data — reps must bring current Bonusly pricing.
- "Rivally has EU data residency." It's real and GA since 2026-07-01 (S15), and was pitched into deals as early as Feb (S05). Counter from data: EMEA rewards catalog thinner than US (S14) and no SCIM (S10) hit EU enterprise accounts hardest. Gap: no source here on Bonusly's own residency status — verify before countering.
- "Rivally is faster to adopt." True on setup (S04). Counter with post-adoption admin cost: no SCIM (S10), no bulk editing (S24), admin tooling lags (S16).
- "Rivally Pulse gives us surveys." It's a separately priced add-on, not bundled (S23, S06) — total cost of the bundle is unverified against their $7 list.
- Prospect fear of painful migration off Rivally (CSV-only exports, S20): needs a Bonusly migration-tools reference to answer — none in the provided data (missing).

## 6. Recent changes
- 2025-11-04: Series C, $40M led by Northgate Ventures (S01, press).
- 2026-03-05: Launched "Rivally Pulse" engagement survey add-on (S06, press).
- 2026-05-09: Hired ex-Workday VP EMEA for European expansion (S11, press).
- 2026-07-01: Dublin office opened; EU data residency GA (S15, press).
- 2026-08-12: Starter list price raised $5 → $7 (S17, pricing page).
- 2026-08-20: Microsoft Teams app v2 in public preview (S19, press).
- 2026-09-01: Pulse exited beta; priced as an add-on, not bundled (S23, press).

## 7. 12-month head-to-head record (deals_with_competitor.csv, Rivally only)
Trailing 12 months = 2025-10 through 2026-09:
- Wins (6): Deal-A9FD43 (2025-10), Deal-7AA785 (2025-11), Deal-44C524 (2025-12), Deal-0D0CD6 (2026-01), Deal-D5B790 (2026-02), Deal-5C636E (2026-03)
- Losses (2): Deal-5645A5 (2026-04), Deal-C6FFAA (2026-05)
- Win rate: 6/8 = 0.75 → 75%.
- Including the 2025-09 loss (Deal-7767F5, just outside the window): 6W / 3L of 9 = 6/9 = 0.667 → 66.7%.
- Trend: five straight wins Oct 2025–Mar 2026, then the two most recent recorded deals both lost (Apr–May 2026).
- Data gaps, stated explicitly:
  - No deals recorded 2026-06 through 2026-09 — unknown whether zero head-to-heads or missing data.
  - S25 (2026-09-03) describes a Bonusly win over Rivally (800 seats, cited analytics depth) that is not in the CSV — the CSV appears stale or incomplete; unreconciled.
  - No win/loss reasons are given in the CSV for any deal except what S25 supplies; the Apr–May losses have no documented loss reason in the data.

## 8. Old-card claims: verification status
- "Pricing starts at $5/user/mo annual (as of 2026-01)" — was sourced (S03, S08), now superseded by $7 (S17). Update to current.
- "Rivally lacks a Slack integration" — CONTRADICTED: Slack integration "worked out of the box" (S04). Remove the claim.
- "Rivally was acquired by WorkHuman in 2025" — UNVERIFIED: no snippet supports it, and it conflicts with Rivally raising its own $40M Series C (S01, 2025-11-04). Strike from the card pending confirmation by CI/legal.
- "Points-based recognition for mid-market" — verified (S02, S04).
- "Strong in EU enterprise with multi-language support" — verified (S12, S11, S15), with the EMEA catalog-thinness caveat (S14).
nurture-sequence-diagnosisPASS 1.00
Plain-English question: What is broken in each email sequence — rates, overlaps, tracking errors?
Prompt — what the model saw
Review the sequences. Per sequence: sent, open rate, reply rate, meeting rate, weakest step. Identify tracking errors (opened above sent), audience overlap between sequences, and the failure mode for anything under 2 percent reply. One change per weak sequence and which to fix first. Under 350 words.
cost $0.1117 · 368s · in 73,854 / out 31,823 tok
✓ tracking_error✓ overlap_pair✓ weakest_step✓ fix_priority

✓ Correct Answer (ground truth)

Plain-English answer: Per-sequence open, reply, and meeting rates with the weakest step (all three weakest at step 3), an opened-above-sent tracking error, and audience overlap flagged.
The answer in one line: fix_priority_first = Cold Outbound - HR Leaders
all_pairwise_overlaps['Cold Outbound - HR Leaders', 'Cold Outbound - People Ops'], ['Expansion Nurture', 'New Logo Nurture']
audience_overlap_pairCold Outbound - HR Leaders, Cold Outbound - People Ops
failure_modeCold Outbound - HR Leaders=opens healthy (~37%) but replies near zero from step 1 -> body/CTA not landing, not a deliverability or subject-line problem
fix_priority_firstCold Outbound - HR Leaders
per_sequence{"New Logo Nurture": {"sent": 1386, "opened": 490, "replied": 90, "meetings": 27, "open_rate": 0.3535, "reply_rate": 0.0649, "meeting_rate": 0.0195, "weakest_step": 3}, "Expansion Nurture": {"sent": 8
tracking_errorsequence=Expansion Nurture; step=2; sent=300; opened=340; issue=opened exceeds sent
under_2pct_replyCold Outbound - HR Leaders
graded checkstracking_error · overlap_pair · weakest_step · fix_priority

Model's actual answer

show the model's full answer
Per-sequence totals [opened+replied+meetings / sent]:

New Logo Nurture: sent 1386 (500+458+428). Open 490/1386=35.4%. Reply 90/1386=6.49%. Meetings 27/1386=1.95%. Weakest: Step 3: reply 18/428=4.21%, meet 6/428=1.40%.

Expansion Nurture: sent 875 (300+300+275). Open 565/875=64.6% inflated. Reply 59/875=6.74%. Meet 12/875=1.37%. Weakest: Step 3: reply 12/275=4.36%, meet 3/275=1.09%.

Cold Outbound - HR Leaders: sent 1785 (600+595+590). Open 545/1785=30.5%. Reply 8/1785=0.45%. Meet 0/1785=0%. Weakest: Step 3: 1/590=0.17% reply, but fails from Step 1: 5/600=0.83%.

Cold Outbound - People Ops: sent 1163 (400+386+377). Open 340/1163=29.2%. Reply 29/1163=2.49%. Meet 6/1163=0.52%. Weakest: Step 3: reply 6/377=1.59%, meet 1/377=0.27%.

Tracking error: Expansion Nurture Step 2: opened 340 > sent 300 = 113.3%. Impossible; double-count or bot/re-open logic. All other steps opened < sent.

Audience overlap: 940 unique keys, 23 in >1 sequence, 0 exact duplicate rows. CT-000301, CT-000624 in both New Logo Nurture + Expansion Nurture. 21 in both Cold Outbound - HR Leaders + Cold Outbound - People Ops: CT-000849, CT-000884, CT-000890, CT-000908, CT-001033, CT-001097, CT-001101, CT-001103, CT-001105, CT-001130, CT-001153, CT-001159, CT-001217, CT-001227, CT-001236, CT-001255, CT-001258, CT-001277, CT-001285, CT-001311, CT-001345.

Failure <2% reply: HR Leaders all steps (0.83%, 0.34%, 0.17%), People Ops Step 3 (1.59%). HR Leaders mode = message/offer mismatch, not deliverability: opens 22.0-40.0% but 0 meetings. People Ops Step 3 mode = sequence fatigue, no new value.

One change: HR Leaders: pause, rewrite Step 1 offer/audience. Expansion: fix Step 2 open tracking before any optimization. People Ops: replace Step 3 with new angle/breakup.

Fix first: Cold Outbound - HR Leaders — largest volume, 0.45% reply, zero meetings.
weekly-marketing-goals-rollupPASS 1.00
Plain-English question: Are we on pace on the quarter's marketing goals?
Prompt — what the model saw
Produce the weekly marketing goals update for the active quarter: SQMs, SQOs, DS2s, closed-lost MIA rate, same-quarter close count, active pipeline coverage against target. For each: QTD actual, target, delta, pace (on/behind/ahead given days elapsed). One paragraph on what moved this week.
cost $0.0028 · 69s · in 488 / out 4,247 tok
✓ metrics_present✓ deltas✓ coverage_ratio✓ pace✓ paragraph

✓ Correct Answer (ground truth)

Plain-English answer: SQMs ahead of pace, SQOs on pace, DS2s behind (35 short of expected at day 66 of 92), with deltas and same-quarter closes.
The answer in one line: days_in_quarter = 92; days_elapsed = 66; coverage_ratio_active_pipeline = 0.75; elapsed_fraction = 0.7174 (+1 more below)
days_in_quarter92
coverage_ratio_active_pipeline0.75
days_elapsed66
elapsed_fraction0.7174
metrics{"SQMs": {"qtd_actual": 230, "target": 300, "delta": -70, "expected_at_day_elapsed": 215.22, "pace": "ahead"}, "SQOs": {"qtd_actual": 84, "target": 120, "delta": -36, "expected_at_day_elapsed": 86.09,
pace_ruleexpected = target * days_elapsed/days_in_quarter; within +/-5% of expected = on, above = ahead, below = behind; lower_better metrics invert.
this_week_paragraph_required1
graded checksmetrics_present · deltas · coverage_ratio · pace · paragraph

Model's actual answer

show the model's full answer
Q3-2026 weekly marketing goals update (from provided files only)

Basis: quarter_meta.csv — Q3-2026 runs 2026-07-01 → 2026-09-30, 92 days, 66 elapsed.
Elapsed fraction = 66 ÷ 92 = 0.7174 → a linear target is "on pace" when actual ≥ 71.7% of target. 26 days remain.

| Metric (alias) | QTD actual | Target | Delta | Pace vs 71.74% |
|---|---|---|---|---|
| SQMs | 230 | 300 | −70 (76.7% of target) | 230 vs 215.2 needed → AHEAD (+14.8) |
| SQOs | 84 | 120 | −36 (70.0% of target) | 84 vs 86.1 needed → BEHIND (−2.1, marginal) |
| DS2s | 40 | 75 | −35 (53.3% of target) | 40 vs 53.8 needed → BEHIND (−13.8) |
| closed_lost_mia | 5 of 25 closed_lost_total = 20.0% rate | 0.1 (lower_better) | +10.0 pts over cap | BEHIND (rate cap; see below) |
| same_quarter_closes | 10 | 20 | −10 (50.0% of target) | 10 vs 14.35 needed → BEHIND (−4.35) |
| active_pipeline | $3,000,000 | $4,000,000 | −$1,000,000 (75.0% of target) | $3.0M vs $2.869M needed → AHEAD (+$130,435) |

Arithmetic detail:
- SQMs: 300 × 0.7174 = 215.2; actual 230 exceeds. Run-rate check: 230/66 = 3.48/day now vs 70/26 = 2.69/day required.
- SQOs: 120 × 0.7174 = 86.1; actual 84 falls 2.1 short. Run-rate: 84/66 = 1.27/day vs 36/26 = 1.39/day required.
- DS2s: 75 × 0.7174 = 53.8; actual 40 short by 13.8. Run-rate: 40/66 = 0.61/day vs 35/26 = 1.35/day required.
- MIA rate: 5 ÷ 25 = 20.0% vs ≤10% cap. Rate is cumulative, not linearly paced; it is over cap now but recoverable — e.g., 0 further MIA with ≥25 more non-MIA closed losses gives 5/50 = 10.0%.
- same_quarter_closes: 20 × 0.7174 = 14.35; actual 10 short by 4.35. Run-rate: 10/66 = 0.152/day vs 10/26 = 0.385/day required.
- active_pipeline: 4,000,000 × 0.7174 = $2,869,565; actual $3.0M exceeds by $130,435. Note this treats a stock metric (pipeline on hand) against a linear-pacing benchmark — coverage is ahead under that convention, but the files give no reason to expect linear accrual.

Derived funnel rates (from provided counts only): SQM→SQO = 84/230 = 36.5%; SQO→DS2 = 40/84 = 47.6%; DS2s vs same-quarter closes = 10/40 = 25.0% of QTD DS2s closed in-quarter.

What moved this week — data limitation stated explicitly: the files contain a single QTD snapshot with no prior-week figures, so week-over-week deltas cannot be computed and I will not invent them. What the snapshot shows structurally: top-of-funnel is carrying the quarter (SQMs ahead of pace, pipeline coverage ahead of the $4.0M target), but conversion is the constraint — 84 SQOs against 230 SQMs (36.5%) and 40 DS2s against 84 SQOs (47.6%) leave DS2s 13.8 behind pace, the largest gap in the set, and only 10 same-quarter closes against a target of 20. Closed-lost MIA rate at 20.0% is double the 10% cap (5 of 25 losses), meaning 20% of lost deals are dying of silence rather than competition. Net read: marketing is generating enough volume; the quarter hinges on advancing existing SQOs into DS2s, pushing in-quarter close dates, and re-engaging MIA-losing deals before the 26 remaining days.
communication 5 tests
ceo-slack-compression0.80
Plain-English question: Compress the forecast analysis into one Slack message the CEO can read in ten seconds.
Prompt — what the model saw
Compress the forecast analysis below into a single Slack message for the CEO. Rules: at most three sentences; under 80 words; lead with the decision or the so-what; keep every distinct number that matters and the one caveat that changes the conclusion; cut any sentence that does not carry a claim, a number, or a directive. No em dashes, no throat-clearing phrases.
cost $0.0016 · 43s · in 111 / out 2,671 tok
✓ max_3_sentences✓ under_80_words✓ weighted_present✗ caveat_present✓ no_em_dash

✓ Correct Answer (ground truth)

Plain-English answer: Three sentences, under 80 words, decision first, every distinct number kept plus the one caveat — all checked by deterministic text rules.
The answer in one line: weighted_number = 115976; must_keep = data-quality caveat — deleting the bad news fails
weighted_number115,976
limitmax 3 sentences, max 80 words, no em-dashes
must_keepdata-quality caveat — deleting the bad news fails
graded checksmax_3_sentences · under_80_words · weighted_present · caveat_present · no_em_dash

Model's actual answer

show the model's full answer
Q3 forecast is $115,977: $44,729 committed across 7 deals plus 35% of $203,565 in best case, 24 deals, $71,248. The risk: 61% of that total rides on best-case deals closing, and 85 of 86 deals have no owner, so rep-level accountability is impossible. Directive: pull committed Deal-D348E1 ($13,770, dated Oct 15) inside Sept 30.
stalled-deal-followup-emailPASS 1.00
Plain-English question: Write a follow-up email for this stalled deal — under 120 words, one prior touch, one ask.
Prompt — what the model saw
Write a follow-up email for this stalled deal. Under 120 words. Reference one specific prior touch from the engagement history, make one clear ask, and match the tone of the prior thread. No invented details.
cost $0.0005 · 14s · in 134 / out 425 tok
✓ word_cap✓ references_touch✓ single_ask

✓ Correct Answer (ground truth)

Plain-English answer: An email that references the July 28 milestone-awards demo, makes one clear ask, and invents nothing — Deal-0D2F7A, 31 days since last engagement.
The answer in one line: amount = 5100; word_cap = 120; days_since_last_engagement = 31; deal_alias = Deal-0D2F7A (+4 more below)
amount5,100
days_since_last_engagement31
word_cap120
allowed_facts14 items — e.g. Deal-0D2F7A, $5,100, DS3, Alex Franklin
allowed_newENG-00001, ENG-00002, ENG-00003, ENG-00004, ENG-00005
deal_aliasDeal-0D2F7A
must_referenceengagement_id=ENG-00002; date=2026-07-28; type=meeting; why=most substantive prior touch; concrete prospect reactions to milestone awards and points catalog
ownerAlex Franklin
single_cta1
stageDS3
tonewarm, concise, professional (matches prior thread)
graded checksword_cap · references_touch · single_ask

Model's actual answer

show the model's full answer
Subject: Re: Pricing for 150 seats — Deal-0D2F7A

Hi,

Following up on the recap I emailed on August 5 with pricing for the 150 seats, after the July 28 demo with your People team. You'd reacted strongly to the automated milestone awards and the points catalog — wanted to check whether anything's shifted since then.

Totally understand if the timing isn't right; recognition budget decisions don't always move on a sales calendar. Could you just let me know where this sits on your side — still a go for a rollout before your close date in late November, or should we park it and reconnect next quarter?

Happy to jump back on a short call if that's easier.

Thanks,
Alex
funnel-digest-paragraphPASS 1.00
Plain-English question: Write the weekly GTM digest — one paragraph per team, every number traceable.
Prompt — what the model saw
Write the weekly GTM digest: one paragraph per team (Marketing, Sales, CS, PLG) in Bonusly brand voice. Every number must trace to the extracts. No customer contact names, no employee performance commentary, no compensation. Four paragraphs, under 400 words total.
cost $0.0046 · 46s · in 12,186 / out 1,553 tok
✓ four_paragraphs✓ word_cap✓ teams

✓ Correct Answer (ground truth)

Plain-English answer: Four paragraphs (Marketing, Sales, CS, PLG) using only the allowed numbers, no contact names, no performance commentary, under 400 words.
The answer in one line: word_cap = 400
word_cap400
allowed_numbers16 items — e.g. 2, 3, 9, 08
bannedcustomer contact names, employee performance commentary, compensation
required_paragraphsMarketing, Sales, CS, PLG
graded checksfour_paragraphs · word_cap · teams

Model's actual answer

show the model's full answer
Weekly GTM digest, week of 2026-08-31

Marketing came in at 46 SQMs against a target of 52, a gap of 6, or 88% of target (46 ÷ 52 = 0.885). The webinar channel carried its weight, delivering 18 of the 46, which is 39% of all SQMs (18 ÷ 46 = 0.391). Cost per SQM on paid search held steady at $150, no drift week over week. The channel mix is doing real work, and closing the 6-SQM gap next week is the obvious play.

Sales turned 14 SQOs into action, setting 9 DS2 meetings along the way. That's 9 DS2 meetings from 14 SQOs, a 64% set rate (9 ÷ 14 = 0.643). New pipeline created hit $310,000 for the week, and 3 deals closed in the same quarter, the fastest path from first meeting to booked revenue. At $310,000 of new pipeline against 3 same-quarter closes, the engine is filling and converting at the same time.

CS saved 2 renewals this week, keeping revenue that was genuinely at risk on the books. Team NPS moved to 61, a real signal that the recognition flywheel is landing with customers. The watch item: 3 open red-flag accounts heading into next week. Two saves is worth celebrating, and those 3 accounts are where next week's attention goes.

PLG added 412 new signups, and activation came in at 31%, which works out to roughly 128 activated accounts (412 × 0.31 = 127.7). The brighter number: 38 companies hit the aha moment of 10 recognition gives, the moment the product starts proving itself without anyone from Bonusly in the room. That's 38 future customers discovering why recognition spreads, one thank-you at a time.

Arithmetic trace: all base figures (46, 52, 18, $150, 14, 9, $310,000, 3, 2, 61, 3, 412, 31%, 38) come directly from the week_extracts.md file for 2026-08-31. Derived ratios (88%, 39%, 64%, ~128) are computed above from those figures only. No data was missing; nothing outside the extracts was used.
partner-digest-section0.67
Plain-English question: Write the partner digest so every partner in the roster appears exactly once.
Prompt — what the model saw
Write the partner digest for the period. Every partner in the roster appears once: active plays get status with numbers, quiet partners get one line. Partner-sourced pipeline count and amount must match the deal data. Nothing omitted, nothing added.
cost $0.0008 · 18s · in 179 / out 950 tok
✓ all_partners✗ pipeline_amount✓ quiet_one_line

✓ Correct Answer (ground truth)

Plain-English answer: Four active partners with status and numbers, four quiet partners with one line each, and partner-sourced pipeline that reconciles to the deal data.
The answer in one line: rule = every roster partner named exactly once; quiet partners get one line; no partners outside roster.
active_no_dealsWorkWell Group
partner_sourced_pipelinecount=5; amount=550000
per_partner_pipeline{"Apex Rewards Co": {"count": 2, "amount": 275000}, "HRCloud Partners": {"count": 1, "amount": 140000}, "CultureBridge": {"count": 2, "amount": 135000}}
quiet_partnersRecogniq, KudosWave, PeopleFirst Advisors, TotalPerk
rosterpartner_name=Apex Rewards Co; status=active | partner_name=HRCloud Partners; status=active | partner_name=CultureBridge; status=active | partner_name=WorkWell Group; status=active | partner_name=Recogniq; status=quiet | partner_name=KudosWave; status=quiet … (+2 more)
ruleevery roster partner named exactly once; quiet partners get one line; no partners outside roster.
graded checksall_partners · pipeline_amount · quiet_one_line

Model's actual answer

show the model's full answer
PARTNER DIGEST — 2026-08-24 to 2026-09-04

ACTIVE PARTNERS

Apex Rewards Co — 2 partner-sourced opps this period, both logged with UTM Source = Partner, both in DS1. Co-webinar locked for 09-15 (Dana Mercer, 2026-08-25). Amount: not in data.

HRCloud Partners — 1 partner-sourced opp this period; integration referral closed the security review and moved to DS2 (Alex Franklin, 2026-08-27). Amount: not in data.

CultureBridge — 2 partner-sourced opps from a lunch-and-learn, both described as "early stage" — stage not specified beyond that, no DS assignment or UTM confirmation given (Farid Osman, 2026-08-29). Amount: not in data.

WorkWell Group — listed active in the roster, but no activity for them appears in the Slack export or deal data for this period. No opps to report; source data missing.

QUIET PARTNERS

Recogniq — quiet; no activity in the period.
KudosWave — quiet; no activity in the period.
PeopleFirst Advisors — quiet; no activity in the period.
TotalPerk — quiet; no activity in the period.

PIPELINE TOTALS (from deal data provided)

Partner-sourced opp count: 2 (Apex Rewards Co) + 1 (HRCloud Partners) + 2 (CultureBridge) = 5 total.
Stage detail: DS1 = 2 (Apex Rewards Co), DS2 = 1 (HRCloud Partners), early stage unspecified = 2 (CultureBridge).

Partner-sourced pipeline AMOUNT: cannot be computed. No dollar amounts exist in any of the provided files (partner_roster.csv, partner_slack.md, or deal data). Per instruction not to invent numbers, the amount is reported as missing rather than estimated.

Coverage check: roster has 8 partners (4 active, 4 quiet); all 8 appear exactly once above; nothing added beyond roster and Slack export.
rfp-answer-from-policyPASS 1.00
Plain-English question: Answer the 10 RFP items from the policy excerpts only — no guessing.
Prompt — what the model saw
Answer the 10 questionnaire items from the policy excerpts only. Cite the excerpt id for every answer. Where the excerpts do not answer the question, mark it unanswerable and say what document would; do not guess. Tone: vendor response, direct.
cost $0.0008 · 20s · in 156 / out 899 tok
✓ answered_count✓ unanswerable_flagged✓ citations✓ no_guess

✓ Correct Answer (ground truth)

Plain-English answer: Seven answered with excerpt citations (Q1-Q7), three flagged unanswerable with the document that would answer them (Q8-Q10).
The answer in one line: answered_count = 7; rule = unanswerable items must be flagged and the document that would answer them named; do not guess.
answered_count7
answer_map{"Q1": ["P1"], "Q2": ["P1"], "Q3": ["P2"], "Q4": ["P3"], "Q5": ["P4"], "Q6": ["P5"], "Q7": ["P6"]}
ruleunanswerable items must be flagged and the document that would answer them named; do not guess.
unanswerableQ8, Q9, Q10
valid_citation_idsP1, P2, P3, P4, P5, P6, P7, P8
graded checksanswered_count · unanswerable_flagged · citations · no_guess

Model's actual answer

show the model's full answer
Q1 — Yes. Customer data is encrypted at rest using AES-256. [P1]

Q2 — Yes. Data in transit is encrypted with TLS 1.2 or higher. [P1]

Q3 — Yes. We maintain a SOC 2 Type II report, renewed annually. It is available to prospects under NDA. [P2]

Q4 — 90 days past the end of the contract. Data is retained for the duration of the contract plus 90 days, after which it is deleted from production systems. [P3]

Q5 — The current subprocessor list is published at trust.bonusly.com. Customers receive 30 days' notice before a new subprocessor is added. [P4]

Q6 — Yes. SAML 2.0 single sign-on is available on the Enterprise plan (SCIM user provisioning is also available on that plan). [P5]

Q7 — Production databases are backed up daily. Backups are encrypted and retained for 35 days. [P6]

Q8 — Unanswerable from the provided excerpts. None of P1–P8 mention ISO 27001. This would be answered by a current ISO 27001 certificate (or a statement of non-certification) from our compliance/security certifications page — we hold a SOC 2 Type II [P2], but that excerpt does not claim ISO 27001 and we will not represent it as such.

Q9 — Unanswerable from the provided excerpts. No uptime or SLA percentage appears in P1–P8. This would be answered by the Master Services Agreement / Service Level Agreement document (or a dedicated SLA page in the trust portal).

Q10 — Unanswerable from the provided excerpts. P1–P8 contain no reference to HIPAA or Business Associate Agreements. This would be answered by our HIPAA/BAA terms or legal agreement addenda (e.g., a BAA offering documented in the trust portal or contract exhibits).

Summary: 7 of 10 items answered directly from policy excerpts (Q1–Q7); Q8, Q9, and Q10 are not covered by the provided excerpts and require the ISO certification record, the MSA/SLA document, and the HIPAA/BAA legal terms respectively.
ops-maintenance 5 tests
skill-file-reconciliationPASS 1.00
Plain-English question: Audit a skill set against its manifest and flag every conflict with a severity and an action.
Prompt — what the model saw
Reconcile the skill set below against its manifest. Find and report (with severity CRITICAL/WARNING/INFO and an action type TRIM_DESC, UPDATE_BODY, DELETE_SKILL, MERGE, or REVIEW): (1) any two skills whose ALWAYS-trigger phrases overlap/duplicate, (2) any circular delegation chain between skills (name it), (3) any delegation target that does not exist (dangling, name it), (4) any version conflict and which skill should survive, (5) how many manifest descriptions exceed 1,024 characters, (6) any hardcoded page ids, dates, or person names in skill bodies, (7) manifest drift in BOTH directions (files with no manifest row, and manifest rows with no file). Make one proposal per finding; rewrite nothing.
cost $0.0115 · 292s · in 121 / out 21,508 tok
✓ duplicate_trigger_found✓ circular_chain✓ dangling_target✓ oversized_desc_count✓ hardcoded_values✓ manifest_drift✓ severity_and_action

✓ Correct Answer (ground truth)

Plain-English answer: A severity-ranked list of TRIM_DESC / UPDATE_BODY / DELETE_SKILL / MERGE / REVIEW actions covering every rule violation found.
The answer in one line: files_audited = 14; rank = CRITICAL / WARNING / INFO with one fix action per issue
files_audited14
findingsduplicate trigger, circular delegation chain, dangling file reference, oversized description, hardcoded values, manifest drift
rankCRITICAL / WARNING / INFO with one fix action per issue
graded checksduplicate_trigger_found · circular_chain · dangling_target · oversized_desc_count · hardcoded_values · manifest_drift · severity_and_action

Model's actual answer

show the model's full answer
GT/SKILL MANIFEST RECONCILIATION — 14 skill files vs skill_manifest.csv. Severity + action per finding, one proposal each. Method notes at end.

────────────────────────────
(1) OVERLAPPING "ALWAYS TRIGGER" PHRASES
────────────────────────────

F1. [CRITICAL | MERGE] email-drafter ⇄ comms-drafter.
Near-duplicate triggers verbatim in both descriptions: "write me an email," "draft a follow-up," "what should I say," "bump email," "contract nudge," pasted-message-review. Coverage lists are near-identical (outbound, follow-ups, post-demo recaps, pricing/contract, EOQ push, renewal/expansion, QBR follow-up, onboarding), the review protocol is identical (rate 1–10, 2–3 rewrites), and the same contract-follow-up benchmark email appears verbatim in both ("Thanks for the update. This is really helpful..."). comms-drafter is a superset (adds support/Intercom/partner lanes); both carry the identical deal-strategy-coach lane marker.
→ Proposal: merge comms-drafter's non-email lanes into email-drafter (or vice versa) as one skill with a channel-router; retire the loser.

F2. [CRITICAL | REVIEW] pipeline-intelligence-report ⇄ weekly-pipeline-report.
Near-identical ALWAYS phrases: "run the pipeline report"/"generate the pipeline report"/"do the pipeline report"; "pipeline update"/"run the pipeline update"/"update the pipeline"; "what's the pipeline look like"/"what does pipeline look like." Neither declares a lane marker against the other, and pipeline-intelligence-report claims "Master pipeline scoring skill — never answer pipeline questions inline without running it," which also collides with next-to-close's conversational shortlist lane (next-to-close disambiguates vs pipeline-intelligence-report, but nothing disambiguates pipeline-intelligence-report vs weekly-pipeline-report).
→ Proposal: add explicit routing lines (SQM/SQO/bookings cadence → weekly-pipeline-report; deal scoring/tiers → pipeline-intelligence-report; shortlist → next-to-close) to both descriptions.

F3. [WARNING | UPDATE_BODY] next-to-close ⇄ deal-strategy-coach.
next-to-close: ALWAYS trigger "which deals are most likely to close." deal-strategy-coach: trigger when a manager/VP "asks which deals are likely to close." Same phrase, no lane marker in either direction.
→ Proposal: in deal-strategy-coach, scope the manager-prep trigger to coaching output and point deal shortlists to next-to-close.

F4. [WARNING | REVIEW] stale-pipeline-report ⇄ deal-strategy-coach.
stale-pipeline-report: "stale deals," "ghost deals," "which deals they haven't touched recently." deal-strategy-coach: "identifies stalled deals." Stale/stalled/ghost trigger sets overlap on the same underlying population.
→ Proposal: define boundary — bulk hygiene list → stale-pipeline-report; single-deal diagnosis/coaching → deal-strategy-coach — and cross-reference in both.

F5. [WARNING | REVIEW] sales-forecast ⇄ next-to-close.
sales-forecast: "what's our number," "what do we think we're going to close." next-to-close: "what's about to close," "what's closing this week." Deal-level vs quarter-level asks collide on the same wording.
→ Proposal: disambiguate in sales-forecast description (quarter number) vs next-to-close (named-deal shortlist).

F6. [WARNING | REVIEW] model-selection vs every skill in the set.
model-selection claims "ALWAYS run this skill at the start of every task, without exception — before any planning, execution, or skill invocation begins... runs FIRST." That universal-preemption phrase duplicates/conflicts with analysis-validator ("Never skip — even on quick check requests", active for whole session) and with pipeline-intelligence-report / partner-digest / closed-lost-analysis per-task "ALWAYS trigger" claims. Nothing sequences model-selection against them.
→ Proposal: qualify model-selection to "annotates the plan; does not alter other skills' mandatory trigger positions," and state its slot once in each report skill's execution plan.

F7. [INFO | REVIEW] analysis-validator / signalforge-claim-compressor / signalforge-feedback each self-label as "final" ("last thing that runs before any output is published" / "Final style pass" / "absolute final step"). The chain order is declared inside signalforge-feedback and signalforge-claim-compressor but not in analysis-validator.
→ Proposal: add the one-line chain (validator → claim-compressor → feedback) to analysis-validator's §1.

────────────────────────────
(2) CIRCULAR DELEGATION
────────────────────────────

F8. [CRITICAL | REVIEW] Cycle: deal-strategy-coach → email-drafter → deal-strategy-coach.
deal-strategy-coach (manager-email section): "use the `email-drafter` skill which automatically retrieves your Gmail signature." email-drafter (lane marker): strategy → deal-strategy-coach; on a "strategy + draft" ask each skill hands the other half back. comms-drafter feeds the same loop (comms-drafter → deal-strategy-coach → email-drafter). The "suggest deal-strategy-coach" phrasing is a soft guard, not a terminal handoff rule.
→ Proposal: make coach→email-drafter a sub-step (signature retrieval only, control returns), and make email-drafter's coach referral terminal after the draft ships.

No other cycles found: next-to-close → pipeline-intelligence-report → closed-lost-analysis is acyclic (closed-lost-analysis's "called from pipeline-intelligence-report" is inbound-only).

────────────────────────────
(3) DANGLING DELEGATION TARGETS (not in the 14-file manifest)
────────────────────────────

F9. [WARNING | REVIEW] prospect-research-multithreading — invoked as a hard prerequisite by comms-drafter ("invoke... first"), email-drafter ("invoke... in Contact Lookup mode first"), and deal-strategy-coach ("Cross-skill handoff"). No manifest row, no file provided.
→ Proposal: verify it exists in the org library; if so add a manifest row / declared external dependency, else remove the invoke-or-block logic.

F10. [WARNING | REVIEW] bonusly-brand — mandatory Step-0 dependency for comms-drafter ("apply the `bonusly-brand` skill"), email-drafter, sales-forecast ("Always reference `bonusly-brand` skill"), and signalforge-claim-compressor routing target. No manifest row.
→ Proposal: same as F9 — declare or restore.

F11. [WARNING | REVIEW] Eight specialist skills in analysis-validator §12.4 delegation table, all dangling: bonusly-data-questions, bonusly-product-questions, bonusly-business-reporting-questions, bonusly-rewards-questions, bonusly-ppp-questions, bonusly-feature-flag-questions, bonusly-deal-desk-questions, bonusly-datadog-questions. Also analysis-validator §11 and signalforge-feedback ("registered in skill-orchestrator as a terminal step") reference skill-orchestrator — dangling.
→ Proposal: reconcile §12.4 and the orchestrator hook against the org skill library; add rows for whatever exists, re-point what doesn't.

F12. [INFO | REVIEW] Path-only external dependencies with no manifest row: signalforge-reports (/mnt/skills/organization/signalforge-reports/SKILL.md, DESIGN-SYSTEM.md, signalforge.css) — mandatory pre-build reads for pipeline-intelligence-report and weekly-pipeline-report; caveman — referenced by signalforge-claim-compressor ("Relationship to Caveman Skill," forked from "JuliusBrussee/caveman").
→ Proposal: mark these as declared external org-skill dependencies so "dangling vs manifest" and "external-by-design" are distinguishable.

────────────────────────────
(4) VERSION CONFLICTS
────────────────────────────

F13. [CRITICAL | UPDATE_BODY] Gong transcript schema conflict.
closed-lost-analysis's required query uses `t.SNIPPET AS transcript_content` from GONG_TRANSCRIPTS_AGG. analysis-validator v3.6 (G1-D, §13.7) states the only transcript source is TRANSCRIPT with "no fallback," and stale-pipeline-report v1.1 states "`GONG_TRANSCRIPTS_AGG` has exactly two columns: `CONVERSATION_KEY` and `TRANSCRIPT`... There is no `SNIPPET`... will error."
→ Surviving skill/statement: analysis-validator v3.6 (corroborated by stale-pipeline-report). Proposal: replace SNIPPET with TRANSCRIPT in closed-lost-analysis's Source-3 query.

F14. [WARNING | UPDATE_BODY] analysis-validator internal version conflict: frontmatter/header/footer/changelog all say v3.6, but the §7 Validation Trail template emits "Validator: analysis-validator v3.2"; §14 also lists v3.5 rows after the v3.6 row.
→ Survive: v3.6. Proposal: fix the trail template to v3.6 (or make it version-injected) and reorder the changelog.

F15. [WARNING | UPDATE_BODY] sales-forecast changelog v1.1 claims "Quarter-agnostic (Q2 → current quarter throughout)" but the body still reads "1A — HubSpot: Open Q2 Deals" and Tab 6 "Q2 Narrative."
→ Survive: v1.1 intent. Proposal: parameterize Q2 references to Q[N].

F16. [WARNING | REVIEW] AE roster divergence: analysis-validator §12.3 "Core 6 AEs" (includes Hugo Lindqvist 77260721, "Updated May 4, 2026") vs pipeline-intelligence-report Phase 1 "AE owner IDs (verified May 2026)" listing 5 — Hugo Lindqvist omitted. Both dated ~the same; G2-F makes the validator roster the resolution authority.
→ Surviving skill: analysis-validator. Proposal: single roster source referenced by pipeline-intelligence-report, both verify-at-runtime.

F17. [WARNING | UPDATE_BODY] signalforge-feedback internal target conflict: body §4 logs to page 2295136266 under "## Feedback Entries"; Activation Checklist points at Build Log page 2247295002 and "## Feedback Log" section.
→ Survive: §4 target page. Proposal: fix checklist to the feedback-log page and section name.

────────────────────────────
(5) MANIFEST DESCRIPTIONS > 1,024 CHARS
────────────────────────────

F18. [INFO | REVIEW] Count = 0.
Arithmetic: values are 656, 897, 996, 792, 965, 676, 945, 1004, 1006, 962, 1006, 708, 762, 656. Max = 1006 < 1024, so 0 of 14 exceed. Three sit within 20 chars of the limit: pipeline-intelligence-report 1006 (headroom 18), signalforge-claim-compressor 1006 (18), partner-digest 1004 (20).
→ Proposal: freeze headroom on those three (any future edit to their descriptions risks silent truncation); trim before adding text.

────────────────────────────
(6) HARDCODED PAGE IDS, DATES, PERSON NAMES IN BODIES
────────────────────────────

F19. [WARNING | UPDATE_BODY] Drifting person names + owner IDs:
- analysis-validator §12.3 (full roster with HubSpot owner IDs: Bryce Harmon 119337721, Hugo Lindqvist 77260721, Dana Mercer 83155923, Alex Franklin 84342457, Cole Ingram 83155924, Gavin Porter 1520255671, 7 CSMs, Alaina Loori 82535637, Shealagh Coughlin 119069206, Ben Castelli 348210196, Amani Phipps 210200121, John Thomas 78303262, Yasmin Wahid 89062643), §10/G1-K "Escalate to Finance (Manish or Amani)".
- pipeline-intelligence-report Phase 1 AE owner IDs (5 listed).
- partner-digest: "Owner: Amani Phipps," Slack ID <@U03QLMBL7AR>, per-partner contact names (Kelli, Jen Lee, Hani, Bryce, Sara).
- weekly-pipeline-report: named recipient "Ben Lavin" in title, flow, and delivery step.
- stale-pipeline-report: trigger "when any AE or Alaina asks," excluded owner ID 55483190 — while its own Phase 2 mandates "Never hardcode rep names or owner IDs" (internal contradiction).
- deal-strategy-coach: ".edu... routed to Farid for manual qualification," "India (routed to Perseus)."
→ Proposal: move rosters/named recipients to one runtime-resolved reference; strip per-person routing to role-based escalation.

F20. [WARNING | REVIEW] Hardcoded infrastructure IDs (routing constants, mostly deliberate but unverified after May 2026):
Confluence — partner-digest (cloudId 73fe98de-a4a3-4869-9f8a-bb1eeed4cf7f, spaceId 1958248479, folder 2286616609, pages 2286321666, 2265382925, 2236940297, 2237825028, 2239365136, 2238283777); sales-forecast (spaceId 2232811524, parent 2232582148); signalforge-feedback (2295136266, 2234417154, 2247295002); deal-strategy-coach (playbook page 2257879045). Slack — stale-pipeline-report channel C0561C1JCPJ. Sheets — weekly-pipeline-report 1CLZeOsElVDF_LF0ZG_t2nfwvhnZ6bpwqM_nX3WEYzcw and 1ENuaEcCuLjdKhMvp8FK3Ys1ek5Aw9ZuOZhsHJJFoB_k. HubSpot org 1973303 embedded in deal URLs in next-to-close, pipeline-intelligence-report, stale-pipeline-report. Stage IDs 150582536/150582537/150582538/150582539/1175632767 duplicated across four skills (analysis-validator §12.2, next-to-close, pipeline-intelligence-report, stale-pipeline-report).
→ Proposal: one shared "system constants" block with a verify-at-runtime directive (pipeline-intelligence-report already models this; the others don't).

F21. [WARNING | UPDATE_BODY] Stale dated snapshots presented as current:
- model-selection: `last_checked: 2026-05-19` with its own rule "more than 14 days past... run the self-update procedure." Today is 2026-09-25: May 12 + Jun 30 + Jul 31 + Aug 31 + Sep 25 = 129 days stale. Its entire registry (Claude Haiku 4.5 / Sonnet 4.6 / Opus 4.7, prices, context windows) is hardcoded and out of its own freshness policy.
- analysis-validator §8 "Expected ranges (as of May 2026)" (~4.5 months old), G1-J anchors "~452,000" and "~110,097", §5 overclaim example "all 3,200+ customers" (partially self-guarded by "re-verify each session").
- closed-lost-analysis: "In the 30-deal AI-field sample from May 2026: 10 of 10...", intervention stats "17% of losses," "8% stated budget," "field confirmed May 2026," case dates "rep vacation May 4–12," "demo on 4/13."
- weekly-pipeline-report: "Q2 (April 1 – June 30, 2026; total ≈ 64–65)" and "Q1 2026 context (static): $365,152 vs. $475,000 plan (77%)... $2,490,532 vs. $3,288,000" — a report for the current quarter carries expired quarter windows.
- stale-pipeline-report: "Don't query all 97 deals serially" — a pinned deal-count snapshot.
- partner-digest: canonical reference "May 16, 2026 issue," "Pipeline Partner Plays Q2/Q3 2026" titles.
→ Proposal: replace dated snapshots with runtime queries or re-verify timestamps; widen the model registry to any current model set only after the self-update runs.

F22. [INFO | UPDATE_BODY] Illustrative named entities baked into rules/examples (dated but arguably intentional): closed-lost-analysis taxonomy cases — "Softheon," "Estee Lauder," "LIFTOFF," "Nestlé," "Ozinga," "MinIO," "Aurora Innovation," "GCash," "StickerYou," "Ethos Cannabis"; signalforge-claim-compressor examples — "Panopto," "Schneider Downs," "Felix Construction"; signalforge-feedback example titles — "Gavin Porter Rep Diagnostic," "Lowe's Conversation Analysis"; analysis-validator G2-F examples — "Dana Mercer," "Gavin Porter."
→ Proposal: keep as historical evidence but tag them "example from [period]" so they're never mistaken for live pipeline state.

────────────────────────────
(7) MANIFEST DRIFT, BOTH DIRECTIONS
────────────────────────────

F23. [INFO | none] No drift found.
Arithmetic: 14 manifest rows (excluding header) vs 14 provided SKILL.md files; orphan files = 14 − 14 matched = 0; rows with no file = 14 − 14 = 0. On every row, `file` == `declared_name` == frontmatter `name`.
→ Proposal: none needed; record the clean reconciliation. Caveat: description_chars and body_lines were taken from skill_manifest.csv as given — not recomputed from the files.

────────────────────────────
METHOD LIMITS
────────────────────────────
- Delegation resolution scope = the 14-file manifest only; F9–F12 targets may exist in an external org library (several carry /mnt/skills paths), which is why they're tagged REVIEW, not DELETE_SKILL.
- No file was rewritten; every finding carries exactly one proposal.
- Nothing in this report is inferred beyond the provided text; where a skill self-guards ("verify at run time," "re-verify each session"), that is noted rather than scored as a violation.
runbook-from-slack-threadPASS 1.00
Plain-English question: Turn the resolved incident thread into a runbook.
Prompt — what the model saw
Turn the resolved incident thread into a runbook: ordered steps with the exact command or action taken, who ran it, how success was verified, and rollback for each step that changed state. Every step must trace to a message in the thread. Add nothing the thread does not contain; where the thread is unclear, mark the step as needing confirmation.
cost $0.0019 · 35s · in 361 / out 2,218 tok
✓ ordered_steps✓ rollback✓ needs_confirmation✓ sourced

✓ Correct Answer (ground truth)

Plain-English answer: Ordered steps from the thread, each with the exact command, who ran it, how success was verified, and rollback where state changed.
The answer in one line: rule = no steps beyond what the thread contains; every step traces to a message id.
ambiguous_steporder=4; msg=M05; mark=needs-confirmation
ruleno steps beyond what the thread contains; every step traces to a message id.
stepsorder=1; msg=M01; actor=Bryce Harmon; action=acknowledge PagerDuty alert, take IC; state_change=False; verified_by=M01 | order=2; msg=M02; actor=Farid Osman; action=check queue depth; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M02 (48,213 pending) | order=3; msg=M04; actor=Farid Osman; action=pause auto-enqueue via feature flag; command=bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'; state_change=True; rollback=bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'; verified_by=M07 (queue depth falling) | order=4; msg=M05; actor=Elena Sinclair; action=cleared the dead set; state_change=True; needs_confirmation=thread gives no command or method for clearing the dead set | order=5; msg=M06; actor=Bryce Harmon; action=scale reward-worker to 6 replicas; command=kubectl scale deployment/reward-worker --replicas=6; state_change=True; rollback=kubectl scale deployment/reward-worker --replicas=3; verified_by=M07/M08 | order=6; msg=M08; actor=Cole Ingram; action=verify queue drained and error rate baseline; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M08 (returns 0) … (+2 more)
graded checksordered_steps · rollback · needs_confirmation · sourced

Model's actual answer

show the model's full answer
RUNBOOK — Reward-worker queue backlog / Redis timeouts
Source: incident thread #incident-2026-08-29-reward-queue (resolved). Every step traces to a message ID; gaps are marked "NEEDS CONFIRMATION."

Trigger [M01]: PagerDuty alert fired for reward-worker queue depth > 10k.

Step 1 — Acknowledge and take IC
- Action: Bryce Harmon acknowledges the PagerDuty alert and takes IC. No command recorded in thread — NEEDS CONFIRMATION.
- Run by: Bryce Harmon [M01]
- Verified by: Not documented — NEEDS CONFIRMATION.
- State change: No. Rollback: N/A.

Step 2 — Assess queue depth
- Command: `bundle exec rake sidekiq:queue_depth`
- Run by: Farid Osman [M02]
- Result: reward queue at 48,213 pending jobs; normal is under 500.
- State change: No. Rollback: N/A.

Step 3 — Inspect the dead set
- Action: Dead set inspection. Command not recorded in thread — NEEDS CONFIRMATION.
- Run by: Farid Osman [M03]
- Result: 112 jobs, all Redis::TimeoutError from around 13:58.
- State change: No. Rollback: N/A.

Step 4 — Pause enqueue (state-changing)
- Command: `bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'`
- Run by: Farid Osman [M04]
- Verified by: No direct verification of the flag is recorded in the thread — NEEDS CONFIRMATION. (The later queue decline in M07 is not isolated proof of this action.)
- Rollback (documented in thread): `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`

Step 5 — Clear the dead set (state-changing; destructive)
- Action: Elena Sinclair states: "While I was in the console I cleared out the dead set." Exact command not recorded — NEEDS CONFIRMATION.
- Run by: Elena Sinclair [M05]
- Verified by: Not documented — NEEDS CONFIRMATION.
- Rollback: Not documented — NEEDS CONFIRMATION. (This is a state-changing action; do not treat it as an approved repeatable remediation step merely because it occurred during the incident.)

Step 6 — Scale workers up (state-changing)
- Command: `kubectl scale deployment/reward-worker --replicas=6` (was 3)
- Run by: Bryce Harmon [M06]
- Verified by: No direct replica-count verification recorded — NEEDS CONFIRMATION. M07/M08 are later queue observations, not isolated proof of this action's effect.
- Rollback (documented in thread): `kubectl scale deployment/reward-worker --replicas=3`

Step 7 — Monitor queue drain
- Action/Command: Measurement command not recorded in thread — NEEDS CONFIRMATION.
- Run by: Farid Osman [M07]
- Result (reported): queue depth 9,400 and falling ~1,200/min.
- State change: No. Rollback: N/A.

Step 8 — Verify recovery
- Command: `bundle exec rake sidekiq:queue_depth`
- Run by: Cole Ingram [M08]
- Verified result: queue depth returns 0; Cole Ingram reports Datadog error rate back to baseline.
- State change: No. Rollback: N/A.

Step 9 — Re-enable enqueue (state-changing)
- Command: `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`
- Run by: Bryce Harmon [M09]
- Verified by: Bryce Harmon reports 40 new jobs processed cleanly in the next 3 minutes.
- Rollback: Not documented — NEEDS CONFIRMATION. (Disabling enqueue is not recorded in the thread as an approved rollback for this step.)

Step 10 — Scale workers back down (state-changing)
- Command: `kubectl scale deployment/reward-worker --replicas=3`
- Run by: Bryce Harmon [M10]
- Verified by: Queue stable at 0; incident declared resolved by Bryce Harmon.
- Rollback: Not documented — NEEDS CONFIRMATION. (Scaling back to 6 is not recorded as an approved rollback.)

Closure [M10]: Incident resolved at 14:55:00Z by Bryce Harmon.

Notes on fidelity:
- All numbers above are copied from the thread (48,213 / 112 / 9,400 / ~1,200/min / 40 jobs / 0); no new metrics computed.
- Actors exactly as recorded: Bryce Harmon (IC), Farid Osman, Elena Sinclair, Cole Ingram.
- Open confirmations before adopting this runbook: the dead-set inspection and clearing commands [M03, M05], the drain-rate measurement [M07], and rollbacks for Steps 5, 9, and 10.
datadog-incident-timelinePASS 1.00
Plain-English question: Sequence the failure from the logs — what broke first, and what cascaded after it?
Prompt — what the model saw
From the log slice, sequence the failure: first error with timestamp, the cascade in order, the service and job involved, and the Datadog query that would confirm the first error. State what the logs do not show.
cost $0.0024 · 56s · in 453 / out 3,496 tok
✓ first_error_ts✓ service_named✓ cascade_ordered✓ query✓ gaps_stated

✓ Correct Answer (ground truth)

Plain-English answer: A reward-service Redis timeout at 14:01:12Z cascades to sidekiq at 14:01:40 and api-gateway at 14:03:05, with the Datadog query that confirms the first error.
The answer in one line: job_class_involved = RewardGiveJob
cascade_orderservice=reward-service; first_error_at=2026-09-03T14:01:12Z | service=sidekiq; first_error_at=2026-09-03T14:01:40Z | service=api-gateway; first_error_at=2026-09-03T14:03:05Z | service=web-app; first_error_at=2026-09-03T14:03:30Z
clean_servicespostgres
confirming_queryservice:reward-service status:error "@message:Redis::TimeoutError"
first_errortimestamp=2026-09-03T14:01:12Z; service=reward-service; message=Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s
job_class_involvedRewardGiveJob
logs_do_not_showroot cause of the Redis timeout, redis-server's own logs, infrastructure metrics (CPU/memory/network), any deploy or config change preceding the incident, who applied mitigation or when (no actor fields in logs)
graded checksfirst_error_ts · service_named · cascade_ordered · query · gaps_stated

Model's actual answer

show the model's full answer
FAILURE SEQUENCE — 2026-09-03, from datadog_logs.csv + sidekiq_jobs.csv only

FIRST ERROR
2026-09-03T14:01:12Z — reward-service, ERROR:
"Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s"
(All prior lines 13:55:13–13:59:30 are INFO; last healthy reward-service activity: "request completed" 13:57:46, last "job enqueued" 13:59:30.)

CASCADE, IN ORDER
1. 14:01:12Z reward-service ERROR — initial Redis timeout (redis-primary:6379, 5s). Root event.
2. 14:01:20 / 14:01:30 / 14:01:40Z reward-service ERROR ×3 — "retry exhausted for RewardGiveJob".
3. 14:01:40Z sidekiq ERROR — "RewardGiveJob failed: Redis::TimeoutError; retrying in 60s". First worker-queue failure.
4. 14:02:28Z sidekiq ERROR — RewardGiveJob failed again.
5. 14:02:30Z sidekiq WARN — "Queue reward depth above 10,000". Backlog builds (~50s after first job failure).
6. 14:03:05Z api-gateway ERROR — "502 upstream timeout calling reward-service /gives". User-facing surface, ~1m53s after first error (14:03:05 − 14:01:12).
7. 14:03:30Z web-app ERROR — "Give form submission failed: upstream 502". End-user impact.
8. 14:03:31–14:06:52Z — repeating loop: sidekiq RewardGiveJob failures ×4 (14:03:31, 14:04:22, 14:05:26, 14:06:47), api-gateway 502s ×4 (14:03:48, 14:04:13, 14:05:16, 14:06:52), web-app 502 failures ×3 (14:04:45, 14:05:42, 14:06:49).
9. 14:22:10Z reward-service INFO — "Redis connection restored; resuming job processing". Endage of outage: 14:22:10 − 14:01:12 = 21 min 58 s.
10. 14:24:45Z sidekiq INFO — "Queue reward depth below 500". Backlog drained, 2m35s after restore.

Error-line count: 19 ERROR + 1 WARN total — reward-service 4, sidekiq 6, api-gateway 5, web-app 4.
Concurrently, sidekiq_jobs.csv logs 16 failed jobs (14:01:46–14:05:50), all Redis::TimeoutError — 12 RewardGiveJob (J-00001–J-00012) + 4 RecognitionDigestJob (J-00013–J-00016).

SERVICES AND JOB INVOLVED
- Origin: reward-service (Redis dependency at redis-primary:6379)
- Queue: sidekiq (reward queue; 16 job failures recorded, depth >10,000)
- Downstream: api-gateway (502s), web-app (give-form failures)
- Job: RewardGiveJob (the failing job named in the logs); sidekiq_jobs.csv shows RecognitionDigestJob also failed on the same error.

DATADOG QUERY TO CONFIRM THE FIRST ERROR
service:reward-service status:error "Redis::TimeoutError" @message:"Connection to Redis at redis-primary:6379 timed out after 5s"
with time range 2026-09-03 13:55:00 → 14:05:00 UTC — earliest matching event should be 14:01:12Z; a broader fallback (drop the @message facet) is service:reward-service status:error "Redis::TimeoutError".

WHAT THE LOGS DO NOT SHOW
- Root cause of the Redis timeout: nothing from redis-primary itself (no slowlog, memory, failover, network, or host metrics). Why it timed out for ~22 min is unexplained.
- No remediation/fix is recorded — "Redis connection restored" at 14:22:10 has no preceding action; automatic recovery vs. human intervention is unknown.
- RecognitionDigestJob never appears in the log lines; it is only evidenced by sidekiq_jobs.csv. Conversely, only 6 sidekiq failure lines vs. 16 failed jobs — the CSV and log slice don't reconcile; one or both are partial.
- Impact volume: number of failed gives, affected users/customers, request traffic, and true queue peak (only "above 10,000" is given; no exact depth).
- 14:07–14:22 gap: after 14:06:52 there are only postgres "checkpoint complete" INFO lines — no reward-service/api-gateway/web-app logs at all. Whether errors stopped or logging simply isn't in the slice is not shown.
- Post-recovery outcome: no log that any of the 16 failed jobs (J-00001–J-00016) retried successfully; queue "below 500" is not zero, and no completion events appear.
- No trace/request IDs, so web-app 502s cannot be individually tied to specific job failures.
feature-flag-state-summaryPASS 1.00
Plain-English question: Summarize what each feature flag controls and who has it on.
Prompt — what the model saw
Summarize the state of the listed feature flags: what each controls per the code excerpt, which companies or segments have it on, targeting rules, and any flag with no code reference or no targets. Company counts must match the export. Do not describe a flag that is not in the export.
cost $0.0015 · 40s · in 173 / out 2,403 tok
✓ orphans✓ no_targets✓ counts_match✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: Recognition streaks on for 42 beta companies, redeem-flow redesign off with a 12-company target list, plus any flag with no code reference or no targets.
The answer in one line: rule = company counts must equal the export; do not describe flags absent from the export.
code_referencedrecognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, redeem_flow_redesign, analytics_dashboard_v3, ms_teams_app_v2
flags{"recognition_streaks_v2": {"state": "on", "targeting_rules": "segment:beta_companies", "company_count": 42}, "points_budget_guardrails": {"state": "on", "targeting_rules": "all_companies", "company_c
flags_with_no_targetspaused_offboard_cleanup
orphan_flags_no_code_referencelegacy_give_modal, survey_boosters_q3, paused_offboard_cleanup
rulecompany counts must equal the export; do not describe flags absent from the export.
graded checksorphans · no_targets · counts_match · no_invented

Model's actual answer

show the model's full answer
FEATURE FLAG STATE — 9 flags in export, 5 on / 4 off

WHAT EACH FLAG CONTROLS (per code excerpt)
Only 2 of 9 flags appear in the code excerpt. For the other 7 the excerpt has no reference, so what they control is not in the provided data and I am not inferring it from the flag name.

1. recognition_streaks_v2 — ON | segment:beta_companies | 42 companies
   Code: app/models/recognition.rb — when enabled for a company, calls StreakTracker.record(give). Controls recording of recognition streaks.

2. points_budget_guardrails — ON | all_companies | 220 companies
   Code: app/services/budget_service.rb — when enabled, calls BudgetService.new(company).enforce!(give). Controls enforcement of points budgets.

FLAGS IN EXPORT WITH NO CODE REFERENCE (7)
3. slack_dm_nudges — ON | segment:region_na | 87 companies
4. redeem_flow_redesign — OFF | targeted_list | 12 companies
5. analytics_dashboard_v3 — ON | segment:tier_three | 65 companies
6. ms_teams_app_v2 — OFF | targeted_list | 9 companies
7. legacy_give_modal — OFF | segment:legacy_plan | 14 companies
8. survey_boosters_q3 — ON | segment:legacy_plan | 7 companies
9. paused_offboard_cleanup — OFF | no targeting rule listed | 0 companies

TARGETING SUMMARY
- all_companies: points_budget_guardrails (220)
- segment:beta_companies: recognition_streaks_v2 (42)
- segment:region_na: slack_dm_nudges (87)
- segment:tier_three: analytics_dashboard_v3 (65)
- segment:legacy_plan: survey_boosters_q3 ON (7) and legacy_give_modal OFF (14) — same segment, different counts; the export gives no reconciliation
- targeted_list (membership not in export): redeem_flow_redesign OFF (12), ms_teams_app_v2 OFF (9)
- No rule at all: paused_offboard_cleanup

FLAG WITH NO CODE / NO TARGETS
- paused_offboard_cleanup: the only flag with no targeting rule and 0 companies; also no code reference.
- redeem_flow_redesign, ms_teams_app_v2, legacy_give_modal: off but carry company counts (12, 9, 14) — targeting defined, flag state off.

ARITHMETIC CHECK (counts as exported)
- ON flags: 42 + 220 + 87 + 65 + 7 = 421 companies
- OFF flags: 12 + 9 + 14 + 0 = 35 companies
- Total flag-company pairs: 421 + 35 = 456
- These are pairs, not unique companies — segments overlap and per-segment sizes are not in the export. The export's own company universe is not stated; 220 (all_companies) is the largest single value.

DATA GAPS (explicit)
- What flags 3–9 control: no code reference provided.
- Who is in the targeted_list for redeem_flow_redesign and ms_teams_app_v2: not in export.
- Whether the 7 vs 14 discrepancy on segment:legacy_plan is stale data or different sub-populations: not determinable from the export.
nda-triage0.67
Plain-English question: Triage each NDA green, yellow, or red — and quote the clause that decides it.
Prompt — what the model saw
Triage each NDA as GREEN (standard approval), YELLOW (counsel review), or RED (full legal review). Cite the clause that drives each classification by quoting it. Catch embedded non-solicits, non-competes, missing carve-outs, and unusual term or governing law. One line of reasoning per NDA.
cost $0.0011 · 31s · in 733 / out 1,392 tok
✓ green✓ yellow✗ red✓ non_solicit_caught✓ clause_quoted✗ correct_labels

✓ Correct Answer (ground truth)

Plain-English answer: NDA-1 green (standard mutual terms), NDA-2 yellow for an embedded 18-month non-solicit, with the driving clause quoted for every classification.
The answer in one line: non_solicit_in = NDA-2
checksclassification_per_nda, clause_quotes_exist_in_document, non_solicit_caught, carveouts_checked
classifications{"NDA-1": {"classification": "GREEN", "driving_reason": "standard mutual NDA: 2-year term, standard carve-outs, Delaware law", "driving_clause_quote": "Carve-outs: information that (a) is or becomes p
non_solicit_inNDA-2
graded checksgreen · yellow · red · non_solicit_caught · clause_quoted · correct_labels

Model's actual answer

show the model's full answer
NDA-1 — GREEN.
Reasoning: mutual evaluation NDA with a complete five-part carve-out set and no embedded restrictions; nothing here exceeds standard terms.
Driving clause (favorable): "Term: two (2) years from the Effective Date; obligations survive for three (3) years after termination." Arithmetic: 2-year term + 3-year survival = up to 5 years total confidentiality exposure from the Effective Date — within normal range. Carve-out (e) "must be disclosed by law or court order" and (d) "independently developed" are both present; governing law "the laws of the State of Delaware, USA" is a mainstream choice; no non-solicit, non-compete, exclusivity, or license grant.

NDA-2 — YELLOW.
Reasoning: the only defect is the carve-out set — counsel review to restore the missing exclusions; no escalation-worthy covenants are embedded.
Driving clause: "Carve-outs: standard — publicly available, previously known, third-party." Despite the word "standard," two carve-outs present in NDA-1 are absent here: (1) no legally-compelled/court-order disclosure exception — without it, complying with a subpoena would technically breach the agreement; (2) no independently-developed exception — our own parallel development is unprotected. Term is "three (3) years" mutual (no survival-after-termination language given), which is unremarkable. Checked and found absent: non-solicit, non-compete, exclusivity.

NDA-3 — CANNOT TRIAGE (data missing).
The file is headed "Three inbound NDAs" but contains only NDA-1 and NDA-2; no text for NDA-3 was provided. I will not invent terms for it — supply the NDA-3 clauses and I'll classify it.