← back to all models ← RevenueBench home
deepseek

deepseek/deepseek-pro-latest

Mean 0.985 · 37/40 perfect tests · $0.99 total · modalities in: text · out: text · each card: the prompt → the correct answer (gold) → this model's actual answer

What do the modalities mean?

Modalities in means what you can send this model: text, images, files, video, or audio. Modalities out means what it can send back. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what this model can handle in real work.

data-crm 5 tests
stage-hygiene-auditPASS 1.00
Plain-English question: How much of the open DS2+ pipeline is missing a substantive why-buys, and is anything missing DS1 timestamps or out of order?
Prompt — what the model saw
Audit stage hygiene on the open pipeline below. Rules: a deal at DS2 or later must have a substantive why-buys (15+ characters, not a placeholder such as tbd, n/a, see notes). A why_buys value of the form LONG(n) means the real text was present but truncated for display; treat it as present and substantive. Every deal should have a DS1 entry timestamp. A deal whose stage entry timestamps are out of order regressed (an earlier-stage timestamp is later than a later-stage timestamp). Report: percent of DS2+ deals non-compliant on why-buys overall and by owner (use the owners table for names), count of deals with no DS1 timestamp, count that regressed, and three process fixes. State definitions before numbers.
cost $0.1381 · 347s · in 108,566 / out 30,813 tok
✓ noncomp_pct✓ no_ds1_0✓ regressed_0✓ per_owner_breakdown✓ definitions_first✓ three_fixes

✓ Correct Answer (ground truth)

Plain-English answer: 81.5% of the 124 DS2+ deals are missing a substantive why-buys — Alex Franklin 48, Cole Ingram 19, Dana Mercer 18, Bryce Harmon 15, Farid Osman 1 — and every deal has a DS1 timestamp with none regressed.
The answer in one line: n_deals = 156; ds2plus = 124; noncomp_pct = 81.5; bryce_noncomp = 15; no_ds1 = 0; regressed = 0 (+1 more below)
n_deals156
bryce_noncomp15
ds2plus124
no_ds10
noncomp_pct81.5
regressed0
per_owner_noncompAlex Franklin=48 of 61 (78.7%); Cole Ingram=19 of 22 (86.4%); Dana Mercer=18 of 22 (81.8%); Bryce Harmon=15 of 18 (83.3%); Farid Osman=1 of 1 (100%)
graded checksnoncomp_pct · no_ds1_0 · regressed_0 · per_owner_breakdown · definitions_first · three_fixes

Model's actual answer

show the model's full answer
Definitions:
- DS2+ = stage in DS2, DS3, DS4, DS5.
- why-buys compliant = 15+ chars substantive, placeholder tbd/n-a/see notes fails, LONG(n) counts as compliant per rule.
- No DS1 = t_ds1 empty.
- Regressed = any earlier-stage timestamp later than a later-stage timestamp.

Counts:
- Total open = 156. DS2+ = 124 = 39 DS2 + 61 DS3 + 14 DS4 + 10 DS5. DS1 = 32.
- why-buys LONG compliant = 23. Non-compliant = 101, all empty, 0 substantive non-LONG text.

why-buys non-compliance DS2+:
- Overall: 101 / 124 = 81.45% non-compliant. Compliant 23 / 124 = 18.55%.
- By owner:
  Bryce Harmon: 15 / 18 = 83.33%
  Farid Osman: 1 / 1 = 100.00%
  Dana Mercer: 18 / 22 = 81.82%
  Cole Ingram: 19 / 22 = 86.36%
  Alex Franklin: 48 / 61 = 78.69%
  Elena Sinclair: 0 DS2+ deals, n/a.
- Check: 15+1+18+19+48 = 101. 18+1+22+22+61 = 124.

Missing DS1: 0 / 156 = 0.

Regressed timestamps: 0 / 156 = 0.

Process fixes:
1. Block DS2+ save/advance when why-buys is empty, under 15 chars, or placeholder.
2. Require why-buys refresh at DS3 and DS4 promotion; audit DS4/DS5 first, highest non-compliance impact.
3. Add entry-criteria checklist on stage change logging t_ds1-t_ds5 automatically; flag out-of-order writes.
owner-resolutionPASS 1.00
Plain-English question: Every open deal has a blank owner name — resolve each one to a real person using the owners table.
Prompt — what the model saw
The deal extract below has an owner name field that is blank on nearly every row and an owner id field that is populated. Using the owners table, resolve each open deal to an owner name. Report: number of deals resolved, deals whose owner id has no match in the owners table (list them), the owner ids that map to archived/deactivated owners (list their names), and the total pipeline amount per resolved owner. Do not guess a name for any unmatched id.
cost $0.0183 · 256s · in 39,020 / out 24,594 tok
✓ resolved_count✓ unmatched_none✓ deactivated_flagged✓ jake_amount✓ justin_amount✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: All 156 deals resolve to six named owners; jake_amount maps to Bryce Harmon ($1,054,144) and justin_amount to Alex Franklin ($624,310).
The answer in one line: jake_amount = 1,054,144; justin_amount = 624,310; resolved = 156
jake_amount1,054,144
justin_amount624,310
resolved156
deactivated
per_ownerBryce Harmon=1,054,144; Alex Franklin=624,310; Dana Mercer=341,195; Cole Ingram=288,161.43; Farid Osman=4,134; Elena Sinclair=2,100
graded checksresolved_count · unmatched_none · deactivated_flagged · jake_amount · justin_amount · no_invented

Model's actual answer

show the model's full answer
Resolved all 156 open deals against the owners table. Verified counts by summing owner_id row counts: 35+67+24+22+7+1 = 156 = total rows.

1. Deals resolved: 156 of 156 (every owner_id in deals_open.csv exists in owners.csv).

2. Owner ids with no match in owners table: NONE (0 deals).

3. Archived/deactivated owners mapped to deals: NONE. The two archived owners in the table (Gavin Porter, 1520255671; Hugo Lindqvist, 77260721) appear in owners.csv but own zero deals in this extract.

4. Total pipeline amount per resolved owner (sum of `amount` across all their open deals):

   Alex Franklin      84342457   67 deals      $624,310.00
   Bryce Harmon      119337721   35 deals    $1,054,144.00
   Dana Mercer        83155923   24 deals      $341,195.00
   Cole Ingram        83155924   22 deals      $288,161.43
   Farid Osman       716654662    7 deals        $4,134.00
   Elena Sinclair    701163055    1 deal         $2,100.00
   ------------------------------------------------------------------
   TOTAL                        156 deals    $2,314,044.43

Arithmetic check on the grand total:
1,054,144.00 + 624,310.00 = 1,678,454.00
1,678,454.00 + 341,195.00 = 2,019,649.00
2,019,649.00 + 288,161.43 = 2,307,810.43
2,307,810.43 + 4,134.00   = 2,311,944.43
2,311,944.43 + 2,100.00   = 2,314,044.43

Notes:
- Cole Ingram's total carries the only fractional amounts in the file (e.g., Deal-AF932D $7,225.40, Deal-FF809F $7,781.20, Deal-A71728 $6,947.50, Deal-42326B $2,480.40, Deal-175395 $4,779.88, Deal-2F3A66 $3,334.80, Deal-D04904 $58,529.25), hence the .43 cents.
- Bryce Harmon's $1,054,144.00 is dominated by two large DS1 PIPELINE deals: Deal-2D1F1B ($240,000) and Deal-66D1FC ($99,000).
- No name was guessed anywhere — all six names come directly from owners.csv.
gong-hubspot-join-integrityPASS 1.00
Plain-English question: How many call records point at deal ids that do not exist in the open pipeline?
Prompt — what the model saw
Check the integrity of the call-to-deal mapping below. The gong table lists calls in the last 90 days keyed to HubSpot deal ids; deals_open lists the current open default-pipeline deals. Report: (a) count of gong rows whose deal id has NO match in deals_open (orphans; list a few sample aliases), (b) duplicate conversation keys (rows where calls_90d exceeds distinct_conversation_keys), (c) among the open deals that ARE at stage DS3 or later, the share that have at least one logged call. Give sample aliases for the orphan finding, all taken from the data. Note: gong_owner_id is a Gong user id, not a HubSpot owner id.
cost $0.0220 · 289s · in 52,012 / out 27,980 tok
✓ orphan_count✓ dup_keys_0✓ ds3_coverage✓ sample_ids_present

✓ Correct Answer (ground truth)

Plain-English answer: 30 Gong call rows are orphans keyed to dead deal ids, none are duplicates, and DS3 call coverage is 29.4%.
The answer in one line: orphans = 30; ds3_cov_pct = 29.4; dup = 0
orphans30
ds3_cov_pct29.4
dup0
orphan_aliasesDeal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42, Deal-36EA09
graded checksorphan_count · dup_keys_0 · ds3_coverage · sample_ids_present

Model's actual answer

show the model's full answer
Integrity check complete. Join was done on hs_deal_id ↔ deal_id (aliases cross-checked; zero alias mismatches on matched rows). Universe: 67 gong rows, 156 open deals.

(a) ORPHANS — 30 of 67 gong rows (30/67 = 44.8%) have no match in deals_open. 37 rows match.

The orphans split into two distinct failure patterns:

Pattern 1 — 25 rows, all with calls_with_ds2_date = 0 AND calls_with_sqo_date = 0, keyed to Gong owners largely unseen in the matched set, and 20+ have last_call dates in Oct–Nov 2026, beyond the 90-day window implied by the rest of the file (which ends ~2026-09-16). Samples: Deal-8FA85D (46 calls), Deal-8FC3F9 (24), Deal-3B7945 (21), Deal-42B265 (21), Deal-9CCC42 (17), Deal-36EA09 (17).

Pattern 2 — 5 rows with normal Gong owners and non-zero ds2/sqo counts, i.e. likely deals that closed/moved off the open default pipeline (or were remapped) but their calls still key to the old id: Deal-3F86A0, Deal-422BA6, Deal-B038F0, Deal-AC944F, Deal-38CA53.

(b) DUPLICATE CONVERSATION KEYS — none. 0 rows have calls_90d > distinct_conversation_keys; every row has calls_90d exactly equal to distinct_conversation_keys (e.g. Deal-C6FE92: 11 = 11; Deal-8FA85D: 46 = 46). No double-counted conversations in this extract.

(c) CALL COVERAGE AT DS3+ — 85 open deals are DS3 or later (DS3 = 61, DS4 = 14, DS5 = 10; 61+14+10 = 85). Of those, 25 have at least one logged call in the gong table.
  Coverage = 25/85 = 29.4%.
Breakdown: DS4/DS5 coverage is 13/24 = 54.2% (e.g. Deal-D348E1, Deal-C26D20, Deal-25F752, Deal-5408B0 have calls); DS3-only coverage is 12/61 = 19.7%. The 60 uncalled DS3+ deals include high-value COMMIT/BEST_CASE rows — Deal-403845 (DS5 COMMIT, $9,000), Deal-9AAE5F (DS4 BEST_CASE, $11,250), Deal-2465CE (DS5 COMMIT, $5,400), Deal-B25F40 (DS3, $40,000), Deal-530B50 (DS3, $31,200) — so either the calls exist but aren't mapped to these deal ids, or they genuinely happened over email/meeting only. The data provided cannot distinguish which.

Note per instructions: gong_owner_id is a Gong user id and was not joined against deals_open.owner_id.
snowflake-sql-trapped-catalogPASS 1.00
Plain-English question: Write one Snowflake query for first-month usage per customer using only the provided, trap-laden catalog.
Prompt — what the model saw
Using ONLY the catalog provided, write one Snowflake SQL query returning, per customer company, for its first calendar month as a customer: unique givers, recognition count, successful redemption count. Read the catalog notes carefully: some tables are stale or unpopulated and must be avoided, and one documented filter (excluding deleted givers) must NOT be applied to historical giving counts. Put a comment above each FROM or JOIN explaining the table choice. Return only fenced SQL.
cost $0.0025 · 71s · in 1,213 / out 4,255 tok
✓ avoids_stale_deals✓ avoids_unpopulated_live✓ uses_cohort_summary✓ succeeded_filter✓ no_giver_deletion_filter✓ returns_sql

✓ Correct Answer (ground truth)

Plain-English answer: A single correct SQL query that uses the right catalog tables and avoids the trap columns — graded on the exact tables and columns.
The answer in one line: use_table = PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
avoid_stalePRODUCTION.HUBSPOT.DEALS (stale)
avoid_unpopulatedV2_LIVE.OBJECTS_DEALS (unpopulated)
forbidden_filterdeleted-giver exclusion must NOT be applied
outputone fenced SELECT with per-FROM/JOIN table-choice comment
required_filtersucceeded redemptions only
use_tablePRODUCTION.PLG.COMPANY_COHORT_SUMMARY
graded checksavoids_stale_deals · avoids_unpopulated_live · uses_cohort_summary · succeeded_filter · no_giver_deletion_filter · returns_sql

Model's actual answer

show the model's full answer
```sql
-- Per-customer-company FIRST CALENDAR MONTH: unique givers, recognitions, successful redemptions.
--
-- DATA GAP (explicit): the catalog contains NO event-level recognition/giving table
-- (no documented giver id, company key, or event timestamp), so none of the three
-- metrics can be computed from raw events. The only first-month, per-company
-- aggregates documented anywhere in the catalog are the M1_* fields on
-- COMPANY_COHORT_SUMMARY. Those are used here.
--
-- SCOPE LIMIT (explicit): COMPANY_COHORT_SUMMARY covers SELF-SERVE companies only.
-- Sales-led companies (HubSpot deals, whose tables are irrelevant to giving anyway)
-- have no giving data in this catalog and are therefore not represented.
--
-- Avoided per catalog notes: HUBSPOT_HUB_1973303.V2_LIVE.OBJECTS_DEALS (unpopulated),
-- PRODUCTION.HUBSPOT.DEALS (stale, 2023-03), PRODUCTION.HUBSPOT.GONG_HUBSPOT_MAP_FAST
-- (retired) — none usable for giving metrics regardless.

SELECT
    -- COMPANY KEY: the catalog excerpt does not document a company-identifier
    -- column for this table. Placeholder used; confirm the real key before running.
    cohort.COMPANY_ID,

    -- "First calendar month AS A CUSTOMER" = calendar month of the first
    -- successful subscription payment (customer = paying), not signup.
    DATE_TRUNC('MONTH', cohort.FIRST_SUB_PAYMENT_DATE) AS first_customer_month,

    -- CAVEAT: M1_USERS is documented as "users", not explicitly "givers"; it is
    -- the closest documented proxy for unique givers in the first month.
    cohort.M1_USERS         AS unique_givers,
    cohort.M1_GIVES         AS recognition_count,
    cohort.M1_REDEMPTIONS   AS successful_redemption_count

-- FROM: the only catalog table with per-company first-month giving/redemption
-- aggregates (M1_USERS, M1_GIVES, M1_REDEMPTIONS) and payment dates; one row
-- per self-serve company.
FROM PRODUCTION.PLG.COMPANY_COHORT_SUMMARY AS cohort

-- NO JOIN to PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2: although it is
-- the documented redemption-event source (STATE = 'succeeded' only), the catalog
-- documents no company key or timestamp for it, so it cannot be scoped per company
-- to the first calendar month. M1_REDEMPTIONS is used instead; per the catalog note,
-- confirm REDEMPTION_RECORDS_V2 alignment before relying on these counts for
-- finance-grade numbers.
--
-- DELIBERATELY NO "exclude deleted givers" filter: per the business rule, it must
-- NOT be applied to historical giving counts (it understates history); the catalog
-- also documents no giver-deletion column to filter on.
--
-- NO JOIN to HubSpot deal/owner/engagement tables: they carry sales-pipeline data,
-- not recognition/redemption events, and offer no join path to giving.

WHERE cohort.FIRST_SUB_PAYMENT_DATE IS NOT NULL   -- only companies that actually
                                                  -- became paying customers
```
crm-completeness-enrichmentPASS 1.00
Plain-English question: Audit the CRM for missing and invalid fields, and propose the fixes.
Prompt — what the model saw
Audit the CRM extract for completeness and propose fixes. Deals need owner, stage, amount, close date, why-buys; companies need industry, employee count, HQ country; contacts need email, title, persona. Report completeness percent per field. Identify duplicate company clusters (name variants, shared domain) with a survivor per cluster. List invalid emails and domain mismatches. Fill missing company fields only where the enrichment export has a matching row; where CRM and enrichment disagree list both and recommend a source. Never invent a value. End with the 10 fixes with the most pipeline amount at stake.
cost $0.0129 · 376s · in 4,393 / out 22,957 tok
✓ dup_clusters✓ invalid_emails✓ survivor_language✓ completeness_pct✓ no_invent

✓ Correct Answer (ground truth)

Plain-English answer: 9 companies missing employee count, 6 missing HQ country, 4 invalid emails, 1 domain mismatch, 2 duplicate clusters, and 25 enrichment matches.
The answer in one line: n_contacts = 52; n_companies = 34; enrichment_matches = 25; invalid_email_count = 4; n_dup_clusters = 2; domain_mismatch_count = 1 (+4 more below)
n_contacts52
domain_mismatch_count1
enrichment_matches25
invalid_email_count4
n_companies34
n_dup_clusters2
company_missingindustry=0; employee_count=9; hq_country=6
domain_mismatchesCT-0011
duplicate_clusters['C-0A092931', 'C-0A092932'], ['C-0A092933', 'C-0A092934']
invalid_emailsCT-0010, CT-0080, CT-0081, CT-0192
graded checksdup_clusters · invalid_emails · survivor_language · completeness_pct · no_invent

Model's actual answer

show the model's full answer
Audit complete. Arithmetic shown inline. One structural caveat up front, flagged honestly:

SCOPE / MISSING DATA
- Extract provided contains only 3 files: companies.csv (34 rows), contacts.csv (51 rows), zoominfo_enrichment.csv (25 rows). No deals extract was provided. Therefore deal fields (owner, stage, amount, close date, why-buys) = 0 of 34... 0 rows available -> completeness cannot be computed, reported as "not in extract."
- Consequence: "10 fixes with the most pipeline amount at stake" cannot be ranked by pipeline $, because no deal amounts exist in the data. I refuse to invent them. Fixes are ranked instead by (a) records corrupted and (b) number of companies that would otherwise stay un-enrichable (contact/domain counts given as the proxy, not dollars).

1. FIELD COMPLETENESS (non-empty / total, % )

COMPANIES (n=34)
- company_alias 34/34 = 100%
- domain 34/34 = 100%
- industry 34/34 = 100% filled, but only 15/34 = 44.1% conform to a clean canonical set (15 rows say "tech"/"Technology" [casing+space variants], 4 say "health care"; 34 - 19 = 15 canonical)
- employee_count 19/34 = 55.9% (15 blank: EC3025, 96039F, 44EA29, D04904, B23205, 60C75F, 425E2A, 7BBDFA, 50D386, 93C8BF, EE9FFB, + the 3 Acme/Globex dup-cluster members and 2 others)
- hq_country 26/34 = 76.5% (8 blank: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB, C-44EA29 counted once; recheck blanks = 2d1f1b, d73b89, 44ea29, d04904, 2c60e5, ee9ffb, 332637-no... blanks are rows 1,7,8,10,21,29 -> C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB, C-C9BB20-no) -> exact blank rows: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB, C-93C8BF-no(has Canada). Verified count = 8 blank rows: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB, C-BA969B-no(has US)... Final verified blanks (8): C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB, C-C9BB20-no(UK), C-60C75F-no(United States) -> see corrected list in section 5; blank-hq rows are 6: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB = 6 blanks -> 28/34 = 82.4%.

Correction (recounted cleanly): hq_country filled = 28/34 = 82.4%; employee_count filled = 19/34 = 55.9% (15 blank).

CONTACTS (n=51)
- email 47/51 = 92.2% filled; valid-format 43/51 = 84.3%
- title 38/51 = 74.5% (13 blank)
- persona 43/51 = 84.3% (8 blank)
- domain 51/51 = 100%
- company_alias 51/51 = 100%, all 51 resolve to a company in companies.csv

DEALS: not present in extract (see caveat).

2. DUPLICATE COMPANY CLUSTERS + SURVIVOR
Confirmed (shared domain or clear variant):
- acme-corp.com: C-0A092931 (Technology, 500, US) + C-0A092932 (tech, 510, USA). Survivor: C-0A092931 (canonical industry value). NOTE survivor/duplicate disagree on employee_count (500 vs 510) - verify before merge, do not auto-pick.
- globex.io: C-0A092933 (SaaS, 200, US) + C-0A092934 (Technology, 200, US). Survivor: C-0A092934 (canonical taxonomy "Technology"); C-0A092933's "SaaS" is an off-taxonomy value.
- 7bbdfa.com + 50d386.com: both "health care", both blank emp, both Canada, no shared domain, alias-distinct. Treated as a probable name/industry-variant duplicate. Survivor: C-7BBDFA (earlier row; evidence is circumstantial).
Low-confidence candidates (not merged - different domains, no name evidence, do NOT invent a merge): EC3025 vs 44EA29 (both tech, emp blank; hq USA vs blank); D04904 vs 60C75F (both Technology, emp blank, hq blank vs United States).

3. INVALID EMAILS
- user@ (no domain): CT-0010, CT-0080, CT-0081, CT-0192
- contact-table domain column gives the fix target: CT-0010/CT-0011 -> 66d1fc.com, CT-0080/CT-0081 -> 92d97d.com, CT-0192 -> 425e2a.com. Only safe reconstruction is local-part + the record's own domain (e.g. user0@66d1fc.com) - flagged as inferred, needs verification.
- No empty emails; all 47 present emails parse except the 4 above.

4. DOMAIN MISMATCHES
- CT-0011: email domain "other-domain.com" != company domain 66d1fc.com. Either a data-entry error or a genuine external contact mis-filed under C-66D1FC.
- All other 50 contacts: email domain == record domain == company domain (CT-0010, CT-0080, CT-0081, CT-0192 fail only on validity, not mismatch).
- Cosmetic casing: C-96039F has domain "96039f.com" (lowercase) consistent everywhere - no real mismatch, just mixed-case hex alias convention.

5. ENRICHMENT FILLS (only where domain matches an enrichment row) + CONFLICTS
Fill employee_count from zoominfo (CRM blank, zi present) - 9 companies:
EC3025<-400, 96039F<-400, 44EA29<-400, D04904<-400, B23205<-400, 60C75F<-400, 7BBDFA<-400, 50D386<-400, B97B4E<-1500 (unchanged), still missing after enrichment (6): 2C60E5, EE9FFB, 93C8BF(no zi row), BA969B... corrected: zi has no row at all for these domains, so no fill: 2c60e5 has zi (340, country blank) -> fill employee 340 (so 10 fills); truly unfilled emp_count: ee9ffb, 93c8bf (no zi row), 425e2a has zi 50 -> CRM blank -> fill 50 (11 fills). Final unfilled after enrichment: ee9ffb.com, 93c8bf.com, and zi-blank-on-emp: none else -> 2 companies + 9 no-zi-row domains still unknown: d73b89... (d73b89 has zi 50) -> domains with NO enrichment row (9): d0662e, b25f40, ee9ffb, c9bb20, 93c8bf, acme-corp, globex.io, 2d1f1b(has zi), 7bbdfa(has zi) -> 9 missing zi rows.
Fill hq_country from zi - 1 company: 2d7423.com (USA already present, no) -> actual fill: none new; zi rows that supply a country CRM lacks: 2d7423 has USA in CRM; 7bbdfa/50d386 Canada (CRM has Canada). CRM blank AND zi present: 66d1fc(CRM USA present) ... net: C-EE9FFB/C-2C60E5 zi rows have blank country -> hq_country not fillable for the 6 blanks except C-44EA29? zi 44ea29 country blank. So 0 new country fills; 6 blanks remain unknown.
CONFLICTS (CRM vs zi disagree - listed as both, with recommendation):
- 66d1fc: "tech" vs "Computer Software"; USA vs United States -> recommend zi (granular taxonomy); country = format-only (zi canonical "United States")
- ec3025: Technology/blank/USA vs Computer Software/400/United States -> emp from zi; industry recommend zi; country format-only
- 96039f: blank vs 400 -> no conflict, fill
- 44ea29, d04904: tech/Technology vs Computer Software -> recommend zi, taxonomy differs
- b23205: Healthcare 400 match; US vs United States format-only
- 2d1f1b: 50 vs 50 match, country blank both -> still missing
- 7bbdfa/50d386: "health care" vs "health care" match; emp fill 400
- 60c75f: tech blank United States vs Computer Software 400 United States -> recommend zi emp, country agrees
- 425e2a: "Tech " 50 USA vs Computer Software 50 United States -> emp fill 50, format-only country
- ba969b vs zi ba969b: zi has NO row for ba969b.com -> excluded from conflict set.
Recommendation rule: zi authoritative for employee_count and industry when CRM value is a non-canonical variant (tech/Technology/health care); prefer CRM when zi is blank (d0662e emp 1500, c6fe92, 31ed2a, 77a95a, aa8dda, e51fb7, 63a874, 92d97d, b97b4e are zi==CRM exact: 7 exact matches: c6fe92, 950043, 92d97d, 31ed2a, aa8dda, e51fb7, d0662e).

6. TOP 10 FIXES (ranked by records corrupted / companies un-enrichable - pipeline amount unavailable, no deals data)
1) Merge acme-corp.com dup (C-0A092931 survivor) - resolve 500 vs 510 with source system.
2) Merge globex.io dup (C-0A092934 survivor, drop "SaaS" row).
3) Merge 7bbdfa.com + 50d386.com (C-7BBDFA survivor) + fill emp 400 from zi.
4) Fix 4 invalid emails (CT-0010, CT-0080, CT-0081, CT-0192) using each record's own domain column - mark as inferred, verify.
5) Investigate CT-0011 domain mismatch (user1@other-domain.com vs 66d1fc.com).
6) Normalize industry taxonomy across all 34 (19 non-canonical: 15 tech/Technology variants, 4 "health care") - unblocks zi matching.
7) Normalize hq_country (USA/United States -> United States; UK -> United Kingdom) across 28 filled rows (4 USA, 8 United States, 12 US, 2 UK = 26 country-coded rows + 6 blank) - 26/34 rows touched.
8) Bulk-fill employee_count from zi for the 11 matched blank rows (EC3025, 96039F, 44EA29, D04904, B23205, 60C75F, 7BBDFA, 50D386, 425E2A, 2C60E5, B97B4E no-blank... B97B4E=1500 so drop) -> 10 fills; 2 remain unknown (ee9ffb, 93c8bf - no zi row).
9) Backfill 12 missing persona rows (CT-0000, 0022, 0041, 0060, 0070, 0082, 0092, 0110, 0132, 0162, 0171, 0172, 0180, 0181 - note title->persona mapping unreliable: CTO appears as both "HR admin" and "champion") -> requires human owner input.
10) Backfill 13 missing titles (CT-0000, 0022, 0081, 0092, 0120, 0121, 0122, 0132, 0141, 0162, 0170, 0171... recount blanks = CT-0000, CT-0022, CT-0081, CT-0092, CT-0120, CT-0121, CT-0122, CT-0132, CT-0141, CT-0162, CT-0170, CT-0180-no(has), CT-0181-no(has) -> 11 blank titles) -> owner input; and verify low-confidence dup pairs (EC3025/44EA29, D04904/60C75F) before any merge.
deal-intelligence 5 tests
closed-lost-classificationPASS 1.00
Plain-English question: Why did we lose the last six months of deals — one primary category per deal, and whose side is the reason on?
Prompt — what the model saw
Classify every deal lost in the last 6 months below into one primary category: pricing, competitor, no decision, timing, product gap, champion left, or other. Also assign a side (Bonusly, buyer, unknown). Then summarize: category counts, the side split, the number of deals where the structured closed_lost_tag clearly disagrees with the free-text reason, and the two patterns most worth acting on. Base everything only on the tag and free-text provided.
cost $0.0214 · 505s · in 16,844 / out 38,695 tok
✓ total_90✓ competitor_top✓ category_counts_present✓ side_split✓ two_patterns

✓ Correct Answer (ground truth)

Plain-English answer: Every closed-lost deal gets exactly one category (pricing, competitor, no decision, timing, product gap, champion left, other) plus a side, per the computed classification.
The answer in one line: total_lost_6mo = 90; top_reason = competitor (26)
total_lost_6mo90
categoriespricing, competitor, no decision, timing, product gap, champion left, other
patternstwo most actionable, tag-vs-text disagreements counted
sideBonusly / buyer / unknown per deal
top_reasoncompetitor (26)
graded checkstotal_90 · competitor_top · category_counts_present · side_split · two_patterns

Model's actual answer

show the model's full answer
Classified all 90 rows in closed_lost_6mo.csv (close dates 2026-07-29 → 2026-09-30; total lost value $1,267,945.16). Where the free-text overrides the tag, the row is flagged with ✗ (tag) → category (text).

PER-DEAL CLASSIFICATION (alias | category | side)
```
Deal-DB0AAC   timing        buyer        Deal-7CB44D   no decision   buyer
Deal-F7F635   competitor    unknown      Deal-0F96AA   competitor    unknown
Deal-AC944F   no decision   buyer        Deal-1BCA50   competitor    buyer
Deal-214060   no decision   buyer        Deal-7CC678   competitor    unknown
Deal-91A056   timing        buyer        Deal-FAC17C   no decision   buyer
Deal-29326C   timing        buyer        Deal-242273   competitor    Bonusly
Deal-5DB9B0   other         unknown      Deal-50E5D8   no decision   buyer
Deal-831B7B   timing        buyer        Deal-A2C349   competitor    buyer
Deal-F97C37   competitor    Bonusly      Deal-9F176A   timing        buyer
Deal-13E9CF   no decision   buyer        Deal-7B2236   pricing       Bonusly
Deal-39E25C   timing        buyer        Deal-AFA56C   no decision   buyer
Deal-7ED004   pricing       buyer        Deal-C7156E   competitor    unknown
Deal-21B045   no decision   buyer        Deal-C33D91   pricing       buyer
Deal-B3ABED   timing        buyer   ✗    Deal-9048EB   product gap   Bonusly ✗
Deal-422BA6   competitor    Bonusly      Deal-5E64CE   timing        buyer   ✗
Deal-ED9AE7   other         buyer   ✗    Deal-8A0992   competitor    unknown
Deal-988493   no decision   buyer        Deal-D0C698   competitor    buyer
Deal-381C8C   competitor    unknown ✗   Deal-69CF3D   timing        buyer
Deal-F308CA   no decision   buyer        Deal-ECBF89   timing        buyer
Deal-F1E8A6   competitor    unknown ✗   Deal-3618CC   product gap   Bonusly ✗
Deal-B6AC09   timing        buyer        Deal-EECC02   competitor    unknown
Deal-70F704   no decision   buyer   ✗   Deal-5AD03E   competitor    unknown ✗
Deal-E6E80A   timing        buyer        Deal-D1A623   timing        buyer
Deal-B038F0   timing        buyer        Deal-413C56   no decision   buyer
Deal-4664E1   no decision   buyer        Deal-47F1A1   competitor    buyer
Deal-175756   timing        buyer        Deal-BF2A98   competitor    buyer
Deal-E74A73   no decision   buyer        Deal-2A292B   no decision   buyer
Deal-DDAB52   competitor    Bonusly      Deal-D1AABF   no decision   buyer
Deal-ACE061   competitor    unknown      Deal-FEDBCB   timing        buyer   ✗
Deal-BB78F3   timing        buyer        Deal-1E7DA9   competitor    unknown
Deal-D48E0B   no decision   buyer        Deal-2BBA21   no decision   buyer
Deal-15DA99   timing        buyer        Deal-286F9C   competitor    unknown
Deal-F4AF5D   timing        buyer        Deal-7FBAC6   no decision   buyer
Deal-79B7A1   timing        buyer        Deal-369281   competitor    buyer
Deal-583ADB   no decision   buyer        Deal-386F6E   no decision   buyer
Deal-8E27DA   other         buyer   ✗   Deal-9FCD0D   competitor    buyer
Deal-2D2F8D   competitor    unknown      Deal-55867E   timing        buyer   ✗
Deal-E0441F   no decision   buyer        Deal-DAFB82   pricing       buyer
Deal-D48E0B is dup-safe     —            Deal-2FEDDB   timing        buyer
Deal-64B19A   competitor    unknown      Deal-3F86A0   no decision   buyer
Deal-096750   no decision   buyer        Deal-F325A5   champion left buyer   ✗
Deal-ABD14C   no decision   buyer        Deal-79E61A   no decision   buyer
Deal-8A119B   pricing       buyer        Deal-AE7C4E   no decision   buyer
Deal-DAB4F1   no decision   buyer        Deal-B4B50F   no decision   buyer
Deal-981AD4   product gap   Bonusly      Deal-DC77FE   competitor    Bonusly
Deal-5885B9   no decision   buyer
```
(90 rows; the stray line above is a typo guard — Deal-FEDBCB appears once, classified timing/buyer.)

CATEGORY COUNTS (n=90; counts + value)
```
no decision     30  33.3%   $279,064
competitor      26  28.9%   $406,235
timing          22  24.4%   $267,411
pricing          5   5.6%   $172,450
product gap      3   3.3%   $ 94,245
other            3   3.3%   $ 34,140
champion left    1   1.1%   $ 14,400
               --  -----   ---------
                90 100.0%  $1,267,945
```
Arithmetic: 30+26+22+5+3+3+1 = 90 ✓

SIDE SPLIT
```
Bonusly  9  (10.0%)  — deals where text names a concrete Bonusly short-
                       coming: Deal-F97C37, Deal-422BA6, Deal-DDAB52,
                       Deal-242273, Deal-7B2236, Deal-9048EB, Deal-3618CC,
                       Deal-981AD4, Deal-DC77FE
buyer   66  (73.3%)  — buyer-side cause (budget, pause, MIA, approval, churn)
unknown 15  (16.7%)  — "went another direction"/"Nothing specific provided"
                        with no diagnostic detail
```
9 + 66 + 15 = 90 ✓

TAG vs FREE-TEXT DISAGREEMENTS
Clearly disagreeing: 1 deal — Deal-8E27DA (tag "Feature Request"; text says they bought a swag provider and "didn't want R&R" — a scope/priority decision, not a Bonusly feature gap).
Additionally, 7 deals where the tag is directly contradicted or unsupported by the text: Deal-B3ABED (Timing tag, text leads "MIA-"), Deal-381C8C and Deal-F1E8A6 (Competitor tag with zero competitor evidence — just "not moving forward"), Deal-9048EB (MIA tag, text says "bad fit… multiple feature gaps"), Deal-5E64CE (Not-a-priority tag, text is a lock-in timing issue with Nectar), Deal-3618CC (Lost DM tag, "Wanted Surveys" = product gap), Deal-5AD03E (Competitor tag, "Wanted more defined budget access" = pricing). Deal-ED9AE7 and Deal-F325A5 are softer mismatches (mixed/leadership causes).
Headline answer: 1 clear disagreement; 8 including the directly-contradicted cases; ~13 total where the tag can't be trusted.

TWO PATTERNS WORTH ACTING ON
1. The "lost" bucket is mostly not lost — it's stalled. no decision + timing = 52/90 deals (52/90 = 57.8%) and $279,064 + $267,411 = $546,475 of $1,267,945 (43.1%). A large share carry explicit future dates ("early 2027," "next year," "2028") — Deal-91A056, Deal-831B7B, Deal-E6E80A, Deal-B038F0, Deal-15DA99, Deal-F4AF5D, Deal-175756, Deal-DAFB82, Deal-B3ABED — yet reps also abandon revivable MIA threads (22 MIA-tagged deals, $254,142, many at only intro-stage contact). Action: a dated re-engagement motion (CRM task cadence keyed to the stated reconnect quarter + a second-life sequence on early-stage MIAs) instead of flat closes.
2. Competitive losses are undiagnosed, and where diagnosed the theme is platform breadth, not price. 12 of 26 competitor deals ($252,635 of $406,235) closed with no stated differentiator ("Went another direction," "Nothing specific provided"), so the loss reasons can't be actioned at all. Where a reason exists, it's breadth/integration pull — Rippl (Deal-DDAB52), ADP TotalSource PEO partnership (Deal-422BA6), onsite points-currency spend (Deal-242273), customization/points-as-dollars (Deal-DC77FE), Awardco surveys (Deal-A2C349) — while Deal-DC77FE explicitly says price "wasn't a factor," and pure pricing is only 5 deals / 5.6% of count. A visible secondary cluster is non-comparable pull (Canadian-provider preference ×2, Paylocity/WorkTango/Nectar incumbency). Action: make a mandatory differentiator field on Competitor closes (win/loss intake), and treat the PEO/incumbent and multi-product-breadth losses as partnership/roadmap intel rather than a discount problem.

Data caveats: Deal-5DB9B0 ("Spam.", n_contacts=1) was excluded from cause analysis beyond "other." Side assignment relies only on the text; the 15 "unknown" sides reflect missing detail, not judgment.
pipeline-tieringPASS 1.00
Plain-English question: Score all 156 open deals into exactly one of LOCK / ACTION / BUILD / REVIVE / WATCH / RISKY.
Prompt — what the model saw
Score the open pipeline below on these signals: forecast category, stage, meetings_30d, contact count, and engagement recency. Assign exactly one tier per deal: LOCK, ACTION, BUILD, REVIVE, WATCH, or RISKY (RISKY = the forecast category disagrees with the engagement evidence). Never tier a deal LOCK if it has zero meetings_30d. Do NOT list every deal. Return ONLY a compact JSON object with exactly these keys: - "tier_counts": object mapping each tier to the number of deals in it (must sum to the total deal count), - "tier_examples": object mapping each tier to up to 3 example deal aliases from the data, - "risky_deals": array of the deal aliases you tiered RISKY, - "lock_violations": number of deals tiered LOCK that have zero meetings_30d (should be 0), - "pipeline_shape": one short paragraph on the overall pipeline shape. Note: inbound_emails_30d is 0 for every row due to a data defect, so treat meetings_30d as the inbound signal.
cost $0.0338 · 537s · in 60,265 / out 45,354 tok
✓ json_returned✓ tier_vocabulary✓ tier_counts_sum✓ risky_deals_listed✓ no_lock_violation_claimed

✓ Correct Answer (ground truth)

Plain-English answer: All 156 open deals tiered exactly once across the six tiers using forecast category, stage, meetings, contacts, and recency.
The answer in one line: total = 156
total156
graded checksjson_returned · tier_vocabulary · tier_counts_sum · risky_deals_listed · no_lock_violation_claimed

Model's actual answer

show the model's full answer
```json
{
  "tier_counts": {
    "LOCK": 6,
    "ACTION": 46,
    "BUILD": 22,
    "REVIVE": 6,
    "WATCH": 41,
    "RISKY": 35
  },
  "tier_examples": {
    "LOCK": ["Deal-25F752", "Deal-C26D20", "Deal-D348E1"],
    "ACTION": ["Deal-C6FE92", "Deal-93C8BF", "Deal-EC3025"],
    "BUILD": ["Deal-523604", "Deal-A414F6", "Deal-01E193"],
    "REVIVE": ["Deal-40522D", "Deal-F0EBBB", "Deal-9DDE86"],
    "WATCH": ["Deal-2D1F1B", "Deal-44EA29", "Deal-D04904"],
    "RISKY": ["Deal-7BBDFA", "Deal-E53952", "Deal-F9A3C1"]
  },
  "risky_deals": ["Deal-E53952", "Deal-5408B0", "Deal-9AAE5F", "Deal-547B2B", "Deal-B7EBD1", "Deal-A2B47C", "Deal-2465CE", "Deal-C61CF7", "Deal-62D607", "Deal-584EE5", "Deal-C6D97A", "Deal-7B3B0F", "Deal-F9A08A", "Deal-0660B4", "Deal-FD9F4E", "Deal-BA571A", "Deal-FC22A3", "Deal-7BBDFA", "Deal-60C2C2", "Deal-4A13AD", "Deal-8AD4A5", "Deal-15D24F", "Deal-9D0060", "Deal-690476", "Deal-635B8E", "Deal-ED725A", "Deal-55164C", "Deal-3BA5EA", "Deal-5FDCE4", "Deal-F336B6", "Deal-5EED42", "Deal-BA3DDC", "Deal-7599B8", "Deal-F9A3C1", "Deal-FA32A0"],
  "lock_violations": 0,
  "pipeline_shape": "156 open deals, $2.314M total (6+46+22+6+41+35=156). Only $72.7K (3.1%) is LOCK — 6 late-stage COMMIT/BEST_CASE deals with >=1 meeting in 30d and contact within ~6 days — against $291.7K (12.6%) RISKY, i.e. COMMIT/BEST_CASE rows with zero meetings_30d, the largest being Deal-7BBDFA ($37,440, last touch ~46d stale), Deal-F9A3C1 ($25,000) and Deal-BA3DDC ($23,400): material forecast inflation at DS4/DS5. The bulk is early-to-mid: 46 ACTION deals ($660K, 28.5%) are actively met/last-touch <=4d DS2-DS4, and 41 WATCH ($1.02M, 44%) are mostly DS1/DS2 PIPELINE with no meetings and stale touches ($901.8K of the WATCH dollars sit in DS1+DS2, incl. Deal-2D1F1B $240,000 untouched since 2026-06-16). 22 BUILD ($187K) are DS1 deals already landing meetings, 6 REVIVE ($82.1K) are mid-stage deals gone quiet >14d. Caveat: 4 deals (Deal-3EED2C, Deal-B936FE, Deal-627646, Deal-57FF13) have no row in engagements_by_deal_90d.csv — engagement unverifiable, tiered WATCH; Deal-57FF13 also has no last_contacted_field. Shape: wide, thin-evidence early funnel with a leaky committed tail."
}
```

Method: RISKY = forecast COMMIT or BEST_CASE with meetings_30d = 0; LOCK = DS4/DS5 + COMMIT/BEST_CASE + meetings_30d >= 1 + recency <= 6 days (so 0 lock violations by construction); REVIVE = non-DS1 with recency > 14 days; ACTION = non-DS1 mid-stage with meetings_30d >= 1 or recency <= 4; BUILD = DS1 with meetings_30d >= 1; WATCH = residual (no meetings, stale, or missing engagement row); recency measured against 2026-09-05, meetings_30d used as the inbound signal per the stated data defect.
call-transcript-extractionPASS 1.00
Plain-English question: Extract the CRM write-back fields from each call transcript as JSON.
Prompt — what the model saw
For each transcript, extract CRM write-back fields as JSON: why-buys (prospect statements only), pain points, stakeholders from the speaker list, budget signal (prospect-stated or null), timeline signal, competitor mentioned (only if the prospect raised it), next step (only if explicitly agreed), objections, confidence. Never fill a field from what the rep said or hoped.
cost $0.0037 · 96s · in 2,305 / out 6,387 tok
✓ json_emitted✓ budget_signal✓ null_discipline✓ stakeholders✓ competitor_prospect_only

✓ Correct Answer (ground truth)

Plain-English answer: Per transcript: prospect-only why-buys, pain points, stakeholders, budget and timeline signals, and competitors — e.g. TX-001: '$40k earmarked', live before open enrollment.
The answer in one line: rep = Alex Franklin
allowed_newTX-001, TX-002, TX-003, TX-004, TX-005, TX-006, Achievers, Bucketlist, Nectar
per_transcript{"TX-001": {"why_buys": ["automating anniversary and birthday awards"], "pain_points": ["HR team of three cannot keep up with it manually", "track everything in a spreadsheet, and people slip through
repAlex Franklin
rep_sourced_nulls{"TX-003": ["budget_signal \u2014 only the rep mentioned pricing/budget"], "TX-004": ["next_step \u2014 rep proposed, prospect said 'no promises'"], "TX-006": ["budget_signal \u2014 rep decoy, prospec
graded checksjson_emitted · budget_signal · null_discipline · stakeholders · competitor_prospect_only

Model's actual answer

show the model's full answer
```json
[
  {
    "transcript_id": "TX-001",
    "deal_alias": "Deal-CFE7F4",
    "why_buys": [
      "Automating anniversary and birthday awards — 'our HR team of three cannot keep up with it manually' (Prospect (VP People))"
    ],
    "pain_points": [
      "Manual anniversary/birthday awards processing overwhelmed by an HR team of three (Prospect (VP People))",
      "Tracking everything in a spreadsheet; 'people slip through the cracks' (Prospect (HR Admin))"
    ],
    "stakeholders": [
      "Prospect (VP People)",
      "Prospect (HR Admin)"
    ],
    "budget_signal": "Prospect-stated: 'about $40k earmarked for engagement tools this fiscal year' (Prospect (VP People))",
    "timeline_signal": "Prospect-stated: live before open enrollment in November (Prospect (VP People))",
    "competitor_mentioned": {
      "name": "Achievers",
      "source": "Prospect (VP People) — 'We looked at Achievers last year, but it was too heavy for a team our size'"
    },
    "next_step": {
      "agreed": true,
      "detail": "Security review with prospect's IT lead on September 12 (prospect: 'Yes — let's do the security review on September 12')"
    },
    "objections": [
      "IT sign-off requires SSO and audit logs (Prospect (HR Admin))"
    ],
    "confidence": {
      "level": "high",
      "rationale": "Prospect-stated budget ($40k earmarked), prospect-stated deadline (November), decision-weight stakeholder (VP People) engaged, concrete mutually agreed next step; only gate is IT security requirement."
    }
  },
  {
    "transcript_id": "TX-002",
    "deal_alias": "Deal-70BB30",
    "why_buys": [
      "Tie recognition to retention for the hourly workforce (Prospect (Head of Total Rewards))"
    ],
    "pain_points": [
      "Regretted turnover in hourly workforce is over 30% (Prospect (Head of Total Rewards))"
    ],
    "stakeholders": [
      "Prospect (Head of Total Rewards)",
      "Prospect (CFO)"
    ],
    "budget_signal": "Prospect-stated: '$25k pilot budget for this quarter' approved by Finance (Prospect (CFO))",
    "timeline_signal": "Prospect-stated: decision wanted by end of September (Prospect (CFO))",
    "competitor_mentioned": null,
    "next_step": {
      "agreed": true,
      "detail": "Rep to send pilot agreement; prospect to route it to legal this week (Prospect (CFO): 'Yes — send the pilot agreement and we'll route it to legal this week')"
    },
    "objections": [
      "Workday integration 'has to be rock solid — that's my one condition' (Prospect (CFO))"
    ],
    "confidence": {
      "level": "high",
      "rationale": "Approved budget ($25k pilot), hard decision deadline (end of September), economic buyer (CFO) in the room, legal routing agreed. Prospect stated rep is the first vendor with a real demo, so no competitive alternative was raised; sole condition is Workday integration quality."
    }
  },
  {
    "transcript_id": "TX-003",
    "deal_alias": "Deal-530B50",
    "why_buys": [
      "Make recognition visible across 12 retail locations (Prospect (People Ops Manager))"
    ],
    "pain_points": [
      "Store managers have zero budget autonomy for on-the-spot recognition today (Prospect (People Ops Manager))"
    ],
    "stakeholders": [
      "Prospect (People Ops Manager)"
    ],
    "budget_signal": null,
    "budget_signal_note": "MISSING: the only pricing figure ($8 per employee per month) was stated by the rep (Alex Franklin), not the prospect. No prospect-stated budget.",
    "timeline_signal": "Prospect-stated: 'no rush on our side until Q1' (Prospect (People Ops Manager))",
    "competitor_mentioned": {
      "name": "Bucketlist",
      "source": "Prospect (People Ops Manager) — 'My CEO used Bucketlist at her last company and liked it'"
    },
    "next_step": {
      "agreed": true,
      "detail": "Short call with the prospect's CEO; People Ops Manager to send two times (prospect: 'Yes, let's schedule a call with our CEO — I'll send two times')"
    },
    "objections": [
      "No urgency on the prospect's side until Q1 (Prospect (People Ops Manager))",
      "CEO must be sold first — 'she decides anything people-related' (Prospect (People Ops Manager)); note: CEO is not in the speaker list, so not counted as a stakeholder"
    ],
    "confidence": {
      "level": "low",
      "rationale": "No prospect-stated budget, prospect-declared no urgency until Q1, decision-maker (CEO) unengaged and positively predisposed to a named competitor (Bucketlist). Only positive signal: agreed CEO intro call."
    }
  },
  {
    "transcript_id": "TX-004",
    "deal_alias": "Deal-180D02",
    "why_buys": [
      "Consolidate three separate recognition tools into one (Prospect (VP People))"
    ],
    "pain_points": [
      "Paying for three tools and none of them talk to the HRIS (Prospect (VP People))"
    ],
    "stakeholders": [
      "Prospect (VP People)",
      "Prospect (IT Security Lead)"
    ],
    "budget_signal": "Prospect-stated approval threshold, not a committed budget: 'If it's under $15k annually, I can approve it without going to the board' (Prospect (VP People))",
    "timeline_signal": "Prospect-stated process constraint: procurement cycle runs six to eight weeks minimum (Prospect (IT Security Lead)). No target go-live or decision date was stated.",
    "competitor_mentioned": null,
    "next_step": {
      "agreed": false,
      "detail": "MISSING: no next step was explicitly agreed. Rep proposed a CFO follow-up; Prospect (VP People) replied 'Maybe — I need to check her calendar, no promises.'"
    },
    "objections": [
      "Security review took three months for the last vendor — 'that's my hesitation' (Prospect (IT Security Lead))",
      "Procurement cycle six to eight weeks minimum (Prospect (IT Security Lead))"
    ],
    "confidence": {
      "level": "medium",
      "rationale": "Clear pain and a prospect-stated single-approver path (under $15k), but no prospect-stated timeline for a decision, no committed budget figure, no next step agreed, and security/procurement friction flagged by the IT stakeholder."
    }
  },
  {
    "transcript_id": "TX-005",
    "deal_alias": "Deal-F8767A",
    "why_buys": [
      "Automate service milestones (Prospect (HR Director))",
      "Analytics on recognition equity across departments (Prospect (HR Director))"
    ],
    "pain_points": [
      "Night-shift teams feel invisible — engagement scores run 20 points lower (Prospect (People Ops Coordinator))",
      "Manual service-milestone processing (implied by 'automate service milestones', Prospect (HR Director))"
    ],
    "stakeholders": [
      "Prospect (HR Director)",
      "Prospect (People Ops Coordinator)"
    ],
    "budget_signal": "Prospect-stated: '$12k approved under our engagement line' (Prospect (HR Director))",
    "timeline_signal": "Prospect-stated: running before the January all-hands (Prospect (HR Director))",
    "competitor_mentioned": {
      "name": "Nectar",
      "source": "Prospect (HR Director) — 'We're mid-pilot with Nectar right now, so you'd need to beat that experience'"
    },
    "next_step": {
      "agreed": true,
      "detail": "Rep to present directly to the prospect's exec team on October 2 (prospect: 'Yes — come present to our exec team on October 2')"
    },
    "objections": [
      "Exec team is skeptical after a failed rollout two years ago (Prospect (HR Director))",
      "Must beat the active Nectar pilot experience (Prospect (HR Director))"
    ],
    "confidence": {
      "level": "high",
      "rationale": "Approved budget ($12k), prospect-stated deadline (January all-hands), and a dated exec presentation agreed. Countervailing risk noted: live incumbent pilot (Nectar) and exec skepticism from a prior failed rollout."
    }
  },
  {
    "transcript_id": "TX-006",
    "deal_alias": "Deal-EE195F",
    "why_buys": [
      "Cut the admin time on service awards (Prospect (HR Manager))"
    ],
    "pain_points": [
      "Prospect personally spends five hours a month ordering and shipping plaques (Prospect (HR Manager))"
    ],
    "stakeholders": [
      "Prospect (HR Manager)"
    ],
    "budget_signal": "Prospect-stated qualitative signal, no amount: 'Budget isn't the issue — time is' (Prospect (HR Manager))",
    "timeline_signal": "Prospect-stated: fiscal year resets in February, so a Q1 start is realistic (Prospect (HR Manager))",
    "competitor_mentioned": null,
    "competitor_note": "Prospect stated no other vendors ('Nobody else'). The alternative they are weighing — building internally — is recorded under objections, not as a competitor vendor.",
    "next_step": {
      "agreed": true,
      "detail": "Rep to send a one-page overview; HR Manager to forward it to the COO this week (prospect: 'Yes — send the one-page overview and I'll forward it to our COO this week')"
    },
    "objections": [
      "COO 'usually prefers building things in-house' (Prospect (HR Manager))",
      "Comparison set includes doing it internally rather than buying (Prospect (HR Manager))"
    ],
    "confidence": {
      "level": "medium",
      "rationale": "Quantified, felt pain (5 hrs/month), budget explicitly not a constraint, Q1 window and an agreed forward step — but no budget amount, the economic buyer (COO) is unengaged and predisposed to building in-house, and the sole stakeholder is a single-threaded champion."
    }
  }
]
```
next-to-closePASS 1.00
Plain-English question: Which 3 deals are closest to signature right now?
Prompt — what the model saw
Which 3 deals are closest to signature right now? Use the CRM data and the Slack thread; Slack often has the real status before the CRM does. For each: deal, amount, why it is close, what is left. Under 250 words.
cost $0.0345 · 24s · in 20,454 / out 2,092 tok
✓ exactly_three✓ top3_named✓ slack_first✓ blocked_excluded

✓ Correct Answer (ground truth)

Plain-English answer: Deal-547B2B ($11,200), Deal-B7EBD1 ($9,000), and Deal-403845 ($9,000) — all DS5 COMMIT deals closing September 10-11, with what is left on each.
blocked_excludeddeal_alias=Deal-2465CE; reason=champion left + procurement freeze (Slack #sales-sparring 2026-09-04)
exactly_three1
ranking_ruleSlack signature-imminent deal ranks first; remaining ranked by stage (DS5>DS4), forecast (COMMIT>BEST_CASE>PIPELINE), close_date asc, amount desc
top3rank=1; deal_alias=Deal-547B2B; amount=11200; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11 | rank=2; deal_alias=Deal-B7EBD1; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-10 | rank=3; deal_alias=Deal-403845; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11
graded checksexactly_three · top3_named · slack_first · blocked_excluded

Model's actual answer

show the model's full answer
1. Deal-547B2B — $11,200 — DS5/COMMIT, close 2026-09-11
Why close: Slack Alex 09:12 — redlines clean, signing page out, VP People signing tomorrow.
Left: signature only.

2. Deal-403845 — $9,000 — DS5/COMMIT, close 2026-09-11
Why close: Slack Dana 10:02 — order form with finance, moving fine.
Left: finance sign-off + signature.

3. Deal-B7EBD1 — $9,000 — DS5/COMMIT, close 2026-09-10
Why close: earliest unblocked COMMIT date (2026-09-10 < 2026-09-11); no Slack block.
Left: missing from data — no Slack update, CRM gives stage only.

Excluded: Deal-2465CE — $5,400 DS5/COMMIT — blocked per Dana 09:20: champion left, procurement froze vendors, Q4 at earliest. Do not count as close despite CRM COMMIT.
product-gap-detectionPASS 1.00
Plain-English question: Which prospect complaints are true product gaps versus rollout or enablement issues?
Prompt — what the model saw
Surface true product gaps raised by prospects in active deals. For each candidate: quote the prospect line, classify as real gap, plan gate, or rollout/enablement issue using the product docs excerpt, and name the deal and amount. Only real gaps go in the summary. Rep language does not count as prospect voice.
cost $0.0102 · 22s · in 739 / out 1,500 tok
✓ real_gaps✓ plan_gate_classified✓ enablement_classified✓ gap_summary

✓ Correct Answer (ground truth)

Plain-English answer: Two real gaps (SCIM provisioning, ADP Workforce Now integration) separated from plan gates and enablement issues, each with the prospect quote and deal.
allowed_newTG-001, TG-002, TG-003, TG-004, TG-005, SCIM user provisioning, ADP Workforce Now integration, Custom report builder
enablement_issuestranscript_id=TG-004; deal_alias=Deal-84DBA6; amount=16000; feature=Slack integration recognition adoption; quote=We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.; why_enablement=Slack integration exists on all plans; managers were never trained
excluded_rep_voicetranscript_id=TG-005; line=Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap.; reason=rep voice, not prospect-raised
plan_gatestranscript_id=TG-003; deal_alias=Deal-CFE7F4; amount=32000; feature=Custom report builder; tier_required=Enterprise; quote=I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?
real_gapstranscript_id=TG-001; deal_alias=Deal-EC3025; amount=62000; feature=SCIM user provisioning; quote=We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.; why_real=not listed in product docs at any tier | transcript_id=TG-002; deal_alias=Deal-D0D6B5; amount=23800; feature=ADP Workforce Now integration; quote=Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.; why_real=supported HRIS list is Workday/BambooHR/Gusto/Rippling only
graded checksreal_gaps · plan_gate_classified · enablement_classified · gap_summary

Model's actual answer

show the model's full answer
Candidate review (prospect voice only):

1. Deal-EC3025 - TG-001
Prospect quote: "We need SCIM user provisioning — our security team will not approve any tool that requires manual account management."
Classification: real gap
Reason: product_docs.md states "SCIM user provisioning and ADP Workforce Now integrations are NOT currently listed as supported capabilities."
Amount: not provided in transcripts_gaps.csv or product_docs.md

2. Deal-D0D6B5 - TG-002
Prospect quote: "Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us."
Classification: real gap
Reason: product_docs.md lists HRIS integrations as "Workday, BambooHR, Gusto, Rippling (Pro and above)" and states ADP Workforce Now is NOT currently listed as supported.
Amount: not provided in transcripts_gaps.csv or product_docs.md

3. Deal-CFE7F4 - TG-003
Prospect quote: "I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?"
Classification: plan gate
Reason: product_docs.md lists "Custom report builder" as Enterprise yes, Core/Pro no. Capability exists, gated by tier.
Amount: not provided

4. Deal-84DBA6 - TG-004
Prospect quote: "We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it."
Classification: rollout/enablement issue
Reason: product_docs.md lists "Slack and Microsoft Teams integration" as yes on all plans. Issue described is training/adoption, not missing capability.
Amount: not provided

5. Deal-36C33F - TG-005
Excluded - no prospect-raised gap. Mobile app gap was raised by Alex Franklin: "Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap." Prospect reply: "Good to know. The web version should be fine for our office staff for now." Rep language does not count as prospect voice.

Arithmetic:
Prospect-raised candidates reviewed = 4 (TG-001 through TG-004)
Rep-raised only = 1 excluded (TG-005)
Real gaps = 2, plan gate = 1, rollout/enablement = 1
2 + 1 + 1 = 4

Summary - real gaps only:
- Deal-EC3025 - "We need SCIM user provisioning — our security team will not approve any tool that requires manual account management." - Amount: missing from provided data
- Deal-D0D6B5 - "Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us." - Amount: missing from provided data
Total at-risk amount for real gaps: cannot calculate - deal amounts not provided.
rep-performance 5 tests
stale-pipeline-by-repPASS 1.00
Plain-English question: Which open deals have had no email, call, or meeting in the last 7 days?
Prompt — what the model saw
The snapshot date is 2026-09-05. Using the open deals and the engagements-by-deal table, list every open deal with no logged email, call, or meeting in the last 7 days. Compute recency from the engagements table's last_email/last_call/last_meeting fields, NOT from any date field on the deal (the deal's last_contacted_field is unreliable). For each stale deal give: deal alias, owner name, stage, amount, days since last contact. Group by owner, order by amount descending within each owner. Then state, per owner, how many stale deals they carry and the total stale amount. (Do not write outreach text.)
cost $0.1516 · 631s · in 126,617 / out 65,850 tok
✓ stale_recall_ge_80✓ grouped_by_owner✓ jake_top_amount✓ kevin_stale_amt✓ jenn_stale_amt✓ stale_count

✓ Correct Answer (ground truth)

Plain-English answer: 65 open deals are stale; the largest is Deal-2D1F1B at $240,000, with Bryce Harmon carrying the most stale dollars at $626,243.
The answer in one line: jake_amt = 626,243; jenn_amt = 261,645; kevin_amt = 252,905.03; top_amount = 240,000; stale_count = 65; top_alias = Deal-2D1F1B
jake_amt626,243
jenn_amt261,645
kevin_amt252,905.03
stale_count65
top_amount240,000
top_aliasDeal-2D1F1B
graded checksstale_recall_ge_80 · grouped_by_owner · jake_top_amount · kevin_stale_amt · jenn_stale_amt · stale_count

Model's actual answer

show the model's full answer
Snapshot date: 2026-09-05. Stale = max(last_email, last_call, last_meeting) from engagements-by-deal table is more than 7 days before 2026-09-05 (i.e. on or before 2026-08-28), or no engagement row exists. Deal last_contacted_field ignored per instructions. Days = 2026-09-05 minus latest engagement date.

65 stale deals total. 2 have no engagement row (stated below).

Bryce Harmon - 13 stale deals
Deal-2D1F1B, DS1, 240000, 81 days (latest 2026-06-16)
Deal-66D1FC, DS1, 99000, 16 days (latest 2026-08-20)
Deal-950043, DS1, 70000, 19 days (latest 2026-08-17)
Deal-B23205, DS1, 45000, 16 days (latest 2026-08-20)
Deal-7BBDFA, DS3, 37440, 46 days (latest 2026-07-21)
Deal-332637, DS2, 36000, 9 days (latest 2026-08-27)
Deal-1BEEBF, DS1, 31500, 19 days (latest 2026-08-17)
Deal-C5658B, DS1, 23400, 16 days (latest 2026-08-20)
Deal-40522D, DS3, 21000, 19 days (latest 2026-08-17)
Deal-F0EBBB, DS3, 11400, 24 days (latest 2026-08-12)
Deal-E25A09, DS1, 6000, 9 days (latest 2026-08-27)
Deal-C9C286, DS2, 5502, 9 days (latest 2026-08-27)
Deal-012CB1, DS1, 1, 23 days (latest 2026-08-13)

Dana Mercer - 14 stale deals
Deal-44EA29, DS2, 60000, 10 days (latest 2026-08-26)
Deal-E51FB7, DS2, 43875, 12 days (latest 2026-08-24)
Deal-B42F46, DS1, 27000, 19 days (latest 2026-08-17)
Deal-BA3DDC, DS3, 23400, 15 days (latest 2026-08-21)
Deal-9DDE86, DS2, 20000, 15 days (latest 2026-08-21)
Deal-215CCA, DS3, 18900, 17 days (latest 2026-08-19)
Deal-5EED42, DS3, 16250, 11 days (latest 2026-08-25)
Deal-57887A, DS2, 15000, 8 days (latest 2026-08-28)
Deal-B7EBD1, DS5, 9000, 16 days (latest 2026-08-20)
Deal-3974EB, DS4, 9000, 8 days (latest 2026-08-28)
Deal-F40F04, DS2, 8100, 15 days (latest 2026-08-21)
Deal-87DDD1, DS1, 5000, 19 days (latest 2026-08-17)
Deal-F336B6, DS3, 4200, 15 days (latest 2026-08-21)
Deal-0660B4, DS4, 1920, 16 days (latest 2026-08-20)

Alex Franklin - 19 stale deals
Deal-CC08D1, DS1, 24000, 16 days (latest 2026-08-20)
Deal-E73427, DS3, 18000, 10 days (latest 2026-08-26)
Deal-885F45, DS2, 9300, 12 days (latest 2026-08-24)
Deal-C2FF3C, DS1, 8316, 10 days (latest 2026-08-26)
Deal-3EED2C, DS2, 7200, unknown - no row in engagements table, recency cannot be computed
Deal-0D2F7A, DS3, 5100, 12 days (latest 2026-08-24)
Deal-6C60D4, DS3, 4800, 12 days (latest 2026-08-24)
Deal-13FEBD, DS2, 4680, 12 days (latest 2026-08-24)
Deal-9D0060, DS3, 3840, 12 days (latest 2026-08-24)
Deal-690476, DS2, 3600, 18 days (latest 2026-08-18)
Deal-C6D97A, DS4, 3240, 8 days (latest 2026-08-28)
Deal-EE195F, DS3, 3120, 8 days (latest 2026-08-28)
Deal-278DEC, DS3, 2700, 8 days (latest 2026-08-28)
Deal-635B8E, DS3, 2600, 18 days (latest 2026-08-18)
Deal-6883F3, DS1, 2400, 16 days (latest 2026-08-20)
Deal-4A13AD, DS3, 2160, 26 days (latest 2026-08-10)
Deal-F67D31, DS2, 1800, 8 days (latest 2026-08-28)
Deal-5FDCE4, DS3, 1600, 12 days (latest 2026-08-24)
Deal-BA571A, DS4, 1080, 18 days (latest 2026-08-18)

Cole Ingram - 18 stale deals
Deal-D04904, DS2, 58529.25, 11 days (latest 2026-08-25)
Deal-B25F40, DS3, 40000, 8 days (latest 2026-08-28)
Deal-813836, DS2, 32175, 11 days (latest 2026-08-25)
Deal-1BA595, DS2, 31750, 11 days (latest 2026-08-25)
Deal-CFE1E8, DS3, 18000, 11 days (latest 2026-08-25)
Deal-CD47A6, DS2, 12168, 11 days (latest 2026-08-25)
Deal-627646, DS3, 11193, 11 days (latest 2026-08-25)
Deal-FF809F, DS2, 7781.2, 11 days (latest 2026-08-25)
Deal-AF932D, DS2, 7225.4, 11 days (latest 2026-08-25)
Deal-A71728, DS2, 6947.5, 11 days (latest 2026-08-25)
Deal-8BC9F5, DS2, 5616, 10 days (latest 2026-08-26)
Deal-175395, DS3, 4779.88, 11 days (latest 2026-08-25)
Deal-481E24, DS3, 4140, 10 days (latest 2026-08-26)
Deal-C7F9BF, DS2, 3360, 11 days (latest 2026-08-25)
Deal-2F3A66, DS3, 3334.8, 11 days (latest 2026-08-25)
Deal-342E96, DS2, 2700, 24 days (latest 2026-08-12)
Deal-E568D5, DS3, 1875, 11 days (latest 2026-08-25)
Deal-FD9F4E, DS5, 1330, 10 days (latest 2026-08-26)

Elena Sinclair - 1 stale deal
Deal-57FF13, DS1, 2100, unknown - no row in engagements table, recency cannot be computed

Farid Osman - 0 stale deals

Per-owner totals (arithmetic shown):
Bryce Harmon: 13 stale. 240000+99000+70000+45000+37440+36000+31500+23400+21000+11400+6000+5502+1 = 626243
Dana Mercer: 14 stale. 60000+43875+27000+23400+20000+18900+16250+15000+9000+9000+8100+5000+4200+1920 = 261645
Alex Franklin: 19 stale. 24000+18000+9300+8316+7200+5100+4800+4680+3840+3600+3240+3120+2700+2600+2400+2160+1800+1600+1080 = 109536
Cole Ingram: 18 stale. 58529.25+40000+32175+31750+18000+12168+11193+7781.2+7225.4+6947.5+5616+4779.88+4140+3360+3334.8+2700+1875+1330 = 252905.03
Elena Sinclair: 1 stale. 2100 = 2100
Farid Osman: 0 stale. Total = 0
Combined stale: 65 deals. 626243+261645+109536+252905.03+2100 = 1252429.03
activity-mix-vs-outcomePASS 1.00
Plain-English question: Per rep, what did the last 30 days of emails, calls, and meetings look like next to the deals that advanced?
Prompt — what the model saw
The snapshot date is 2026-09-05. For each sales rep (owner), compute the last-30-day totals of emails, calls, and meetings from the engagements table, and the number of deals that entered DS2 in the last 30 days (t_ds2 within 30 days of the snapshot). Report per rep: the activity mix as percentages (emails/calls/meetings share of that rep's total activities), total activities per DS2 entry, and rank the reps by that efficiency ratio (lowest activities per DS2 entry = most efficient). Name the most efficient rep and the highest-volume rep, and note if they differ. Do not attribute cause.
cost $0.1065 · 392s · in 109,484 / out 34,380 tok
✓ alex_ds2_30d✓ justin_volume_high✓ efficient_named✓ mix_present✓ no_causal_claim

✓ Correct Answer (ground truth)

Plain-English answer: Alex Franklin is the standout signal with 18 deals entering DS2 in 30 days (next closest: Bryce Harmon with 4), alongside per-rep activity totals.
The answer in one line: alex_ds2_30d = 18; bryce_ds2_30d = 4; note = legacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
alex_ds2_30d18
bryce_ds2_30d4
notelegacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
graded checksalex_ds2_30d · justin_volume_high · efficient_named · mix_present · no_causal_claim

Model's actual answer

show the model's full answer
Snapshot: 2026-09-05. Last-30-day window used: 2026-08-06 to 2026-09-05 inclusive (snapshot minus 30 days).
t_ds2 counted if 2026-08-06 <= t_ds2 <= 2026-09-05. Blank t_ds2 = not counted.

Data coverage note:
- All rows in engagements_by_deal_90d.csv match a deal_id in deals_open.csv.
- 2 deals in deals_open.csv have no engagement row, activity missing for those deals: Deal-3EED2C (owner 84342457), Deal-57FF13 (owner 701163055). Activity totals below sum only rows present.
- Owners 1520255671 Gavin Porter (archived) and 77260721 Hugo Lindqvist (archived) have 0 deals in deals_open.csv, so 0 activities, 0 DS2 entries.

Per rep totals (emails_30d + calls_30d + meetings_30d = total):

84342457 Alex Franklin:
 emails 307, calls 36, meetings 41
 total = 307+36+41 = 384
 mix = 307/384=79.9% emails, 36/384=9.4% calls, 41/384=10.7% meetings
 DS2 entries = 18:
  Deal-403845 2026-09-02, Deal-1FC049 2026-09-03, Deal-3EED2C 2026-09-03,
  Deal-7FA0C3 2026-08-07, Deal-E531A6 2026-08-07, Deal-5296C9 2026-08-28,
  Deal-36C33F 2026-08-11, Deal-EE195F 2026-08-06, Deal-F436DA 2026-08-19,
  Deal-317E6F 2026-08-12, Deal-D1E6C2 2026-08-11, Deal-D9A72E 2026-08-06,
  Deal-CA5E44 2026-08-24, Deal-4F775F 2026-08-17, Deal-898FC5 2026-08-28,
  Deal-46988D 2026-08-26, Deal-E73427 2026-08-28, Deal-92D97D 2026-09-02
 activities per DS2 = 384/18 = 21.33

119337721 Bryce Harmon:
 emails 162, calls 0, meetings 43
 total = 162+0+43 = 205
 mix = 162/205=79.0% emails, 0/205=0.0% calls, 43/205=21.0% meetings
 DS2 entries = 4:
  Deal-25F752 2026-08-10, Deal-D73B89 2026-09-03,
  Deal-CA7DC0 2026-08-12, Deal-1CCE5C 2026-08-06
 activities per DS2 = 205/4 = 51.25

83155923 Dana Mercer:
 emails 84, calls 18, meetings 11
 total = 84+18+11 = 113
 mix = 84/113=74.3% emails, 18/113=15.9% calls, 11/113=9.7% meetings
 DS2 entries = 1:
  Deal-57887A 2026-08-07
 activities per DS2 = 113/1 = 113.0

83155924 Cole Ingram:
 emails 96, calls 14, meetings 1
 total = 96+14+1 = 111
 mix = 96/111=86.5% emails, 14/111=12.6% calls, 1/111=0.9% meetings
 DS2 entries = 2:
  Deal-42326B 2026-08-26, Deal-1BA595 2026-08-12
 activities per DS2 = 111/2 = 55.5

716654662 Farid Osman:
 emails 38, calls 0, meetings 34
 total = 38+0+34 = 72
 mix = 38/72=52.8% emails, 0/72=0.0% calls, 34/72=47.2% meetings
 DS2 entries = 1:
  Deal-499BF6 2026-08-26
 activities per DS2 = 72/1 = 72.0

701163055 Elena Sinclair:
 emails 0, calls 0, meetings 0 from available rows; 1 deal Deal-57FF13 has no engagement row, data missing
 total available = 0
 mix = undefined (0 total)
 DS2 entries = 0
 activities per DS2 = undefined (division by zero)

1520255671 Gavin Porter: 0 activities, 0 DS2 entries, ratio undefined.
77260721 Hugo Lindqvist: 0 activities, 0 DS2 entries, ratio undefined.

Efficiency rank (lowest activities per DS2 = rank 1):
1 Alex Franklin 21.33
2 Bryce Harmon 51.25
3 Cole Ingram 55.5
4 Farid Osman 72.0
5 Dana Mercer 113.0
Unranked: Elena Sinclair, Gavin Porter, Hugo Lindqvist (0 DS2 entries, undefined ratio)

Most efficient rep: Alex Franklin at 21.33 activities per DS2 entry (384/18).
Highest-volume rep: Alex Franklin at 384 total activities.
They do not differ; same rep.
ae-qtd-scorecardPASS 1.00
Plain-English question: Build Alex Franklin's QTD scorecard: bookings versus quota, pipeline, and conversion.
Prompt — what the model saw
Build a QTD scorecard for the named AE as of 2026-09-05: bookings vs quota with attainment percent, new vs expansion split, active pipeline by stage with amount, rolling 90-day DS2-to-won rate, win and loss counts with top loss reason, activity volume by type in the last 30 days. Exclude any closed-won deal dated before the quarter. Close with three coaching observations grounded in the numbers.
cost $0.0176 · 309s · in 28,495 / out 26,624 tok
✓ bookings✓ quota✓ attainment✓ new_expansion_split✓ top_loss_reason✓ coaching_obs

✓ Correct Answer (ground truth)

Plain-English answer: 75% attainment ($150,000 on a $200,000 quota), $113,500 new versus $36,500 expansion, pipeline by stage, and the rolling DS2-to-won rate.
The answer in one line: quota = 200000; bookings = 150000; new_bookings = 113500; expansion_bookings = 36500; open_deals_count = 125; attainment_pct = 75; owner = Alex Franklin (+12 more below)
quota200,000
attainment_pct75
bookings150,000
expansion_bookings36,500
expansion_count3
losses_qtd27
new_bookings113,500
new_count5
open_deals_count125
wins_qtd8
activity_30d_totalsemails_30d=807; calls_30d=112; meetings_30d=128; notes_30d=50
allowed_new18 items — e.g. Deal-A1C3E5, Deal-B7D2F4, Deal-C9E1A6, Deal-D4B8C2
ds2_to_wonwindow_start=2026-06-07; won_in_window=8; closed_in_window=35; rate=0.2286
ownerAlex Franklin
pipeline_by_stage{"DS1": {"count": 20, "amount": 284621}, "DS2": {"count": 28, "amount": 353760}, "DS3": {"count": 67, "amount": 552705}, "DS4": {"count": 5, "amount": 23574}, "DS5": {"count": 5, "amount": 45730}}
prior_quarter_excludeddeal_alias=Deal-B3E6F1; close_date=2026-06-20; amount=24000
quarter2026-Q3
snapshot_date2026-09-05
top_loss_reasonreason=Lost- Timing (1 year or more); count=13
graded checksbookings · quota · attainment · new_expansion_split · top_loss_reason · coaching_obs

Model's actual answer

show the model's full answer
AE SCORECARD — Alex Franklin — QTD as of 2026-09-05 (Q3 = Jul 1–Sep 30; data joined on deal_id, 161 deals, 161 engagement rows, 1:1 match)

1. BOOKINGS VS QUOTA
- Eligible closed-won (close_date within quarter, on/before snapshot): 8 deals
  40,000 + 20,000 + 35,000 + 21,000 + 12,000 + 11,000 + 6,500 + 4,500 = $150,000
- Excluded: Deal-B3E6F1 ($24,000, closed 2026-06-20 — before quarter start)
- Quota (quota.csv, 2026-Q3): $200,000
- Attainment: 150,000 / 200,000 = 75.0% | Gap to quota: $50,000

2. NEW VS EXPANSION SPLIT
- New: 5 deals — Deal-A1C3E5, Deal-B7D2F4, Deal-C9E1A6, Deal-D4B8C2, Deal-E6F3A9 = $113,500 (75.7%)
- Expansion: 3 deals — Deal-F2C7D8, Deal-A8B4D6, Deal-C5D9E2 = $36,500 (24.3%)
- Check: 113,500 + 36,500 = 150,000 ✓ | Avg new $22,700 vs avg expansion $12,167

3. ACTIVE PIPELINE BY STAGE (status=open; 125 deals, $1,260,390 total)
- DS1: 20 deals, $284,621
- DS2: 28 deals, $353,760
- DS3: 67 deals, $552,705
- DS4:  5 deals, $23,574
- DS5:  5 deals, $45,730
- Sum check: 284,621+353,760+552,705+23,574+45,730 = 1,260,390 ✓
- Note: Deal-7A2454 (DS3, $1,275) is still open past its 2026-09-04 close date. Only $108,088 of open pipeline (21 deals) has an expected close remaining inside Q3 vs the $50,000 gap.

4. ROLLING 90-DAY DS2-TO-WON RATE
- Window: entered_ds2 in 2026-06-07 → 2026-09-05 (snapshot minus 90 days)
- Cohort: 111 deals entered DS2 | 8 won | 27 lost | 76 still open
- Rate (won / all who entered DS2): 8 / 111 = 7.2%
- Rate (won / decided only): 8 / (8+27) = 8/35 = 22.9%

5. WINS AND LOSSES (closed within quarter, through snapshot)
- Wins: 8 ($150,000) | Losses: 27 ($329,272)
- Win rate on decided deals: 8 / 35 = 22.9%
- Loss reasons by count:
  Lost- Timing (1 year or more): 13 deals, $184,681  ← top reason
  MIA: 5 deals, $45,831
  Competitor: 5 deals, $49,020
  Lost DM: 2 deals, $17,940
  Feature Request: 1 deal, $21,000
  Lost- Does not fit ICP (write in notes): 1 deal, $10,800

6. ACTIVITY VOLUME — LAST 30 DAYS (ae_engagements.csv, all deals)
- Emails: 807 | Calls: 112 | Meetings: 128 | Notes: 50 | Total touchpoints: 1,097
- Per-deal averages: WON deals — 11.1 emails / 3.9 calls / 2.9 meetings / 2.6 notes vs LOST deals — 3.8 / 0.8 / 0.6 / 0.7
- 62 of 125 open deals had zero calls + zero meetings in the last 30 days

COACHING OBSERVATIONS
1. Multi-threading separates wins from losses. Won deals averaged ~5x the calls (3.9 vs 0.8) and ~5x the meetings (2.9 vs 0.6) of lost deals in the last 30 days, while email volume differs only ~3x. Alex is over-indexed on email (807 of 1,097 touches) and under-talking 62 of 125 open deals; converting email volume into calls/meetings is the highest-leverage behavior change.
2. Pipeline is bloated and low-yield. $1.26M open across 125 deals (avg $10,083) against a 22.9% decided win rate and a 7.2% 90-day DS2-to-won rate means only about $91k of the $1.26M would convert at current rates — less than the $50,000 gap plus Q4 needs. The $50,000 gap is realistically closeable from the $108,088 (21 deals) with Q3 close dates, so those should be the priority; 5 DS2 deals ($51,100, e.g. Deal-5BFE3B in DS2 since 2026-01-05) are stale and the timing-loss pattern suggests many open deals will never close.
3. Timing losses are a qualification problem, not a closing problem. 13 of 27 losses ($184,681 — 56% of lost value) cite "Lost- Timing (1 year or more)", and one loss (Deal-5DB9B0) admits "Does not fit ICP." Deals are being advanced to DS2/loss without a confirmed budget/buying-window trigger; tightening upfront discovery on timeline before DS2 would cut the ~90-day dead-weight cohort and free capacity for the meetings that correlate with wins.
multithreading-gapPASS 1.00
Plain-English question: Which open deals are single-threaded or under-threaded?
Prompt — what the model saw
Find every open deal that is single-threaded (fewer than 2 active contacts) or under-threaded (fewer than 3, or all contacts in one persona). Active means engaged in the last 60 days and not marked former. Personas: economic buyer, champion, HR admin, IT security, finance. For each flagged deal: amount, stage, active contact count, personas present, personas missing, the most valuable persona to add given the stage, and any on-file unengaged contact at that company who fits; otherwise say none on file.
cost $0.0041 · 106s · in 2,545 / out 7,181 tok
✓ flagged_recall✓ persona_language✓ single_thread_term✓ on_file_fit

✓ Correct Answer (ground truth)

Plain-English answer: Deals with fewer than 2 active contacts (single-threaded) or fewer than 3, or all in one persona (under-threaded), each flagged with the missing persona.
The answer in one line: active_window_start = 2026-07-07
active_window_start2026-07-07
allowed_new59 items — e.g. C-077A0E, C-0D15DF, C-1BDFE0, C-2AE3AA
best_persona_ruleDS2/DS3: economic buyer > IT security > finance > HR admin > champion; DS4/DS5: economic buyer > finance > IT security > HR admin > champion; first persona not present among active contacts
flag_ruleflag if active_contacts < 2 (single-threaded) or active_contacts < 3 or all active contacts share one persona (under-threaded); active = engaged within 60 days and not former
flaggeddeal_alias=Deal-36C33F; amount=15000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-50D386; amount=36000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5408B0; amount=14850; stage=DS4; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5BFE3B; amount=18000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-885F45; amount=9300; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=IT security | deal_alias=Deal-92D97D; amount=60000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer … (+5 more)
healthy_not_flaggedDeal-4B0BEB, Deal-84DBA6, Deal-D348E1
persona_universeeconomic buyer, champion, HR admin, IT security, finance
graded checksflagged_recall · persona_language · single_thread_term · on_file_fit

Model's actual answer

show the model's full answer
Two fields the task requires are absent from the provided files: deal_contacts.csv and unengaged_contacts.csv contain no amount and no stage columns. So "amount," "stage," and "the most valuable persona to add given the stage" cannot be computed from the data — I flag that per deal and give a stage-independent gap suggestion instead, clearly labeled as not stage-derived. No numbers invented.

Method
- Reference date = 2026-09-25 (today). Active window = last 60 days: 2026-09-25 − 60 = 2026-07-27 cutoff; contact active if last_engaged_date >= 2026-07-27 AND is_former = false.
- Flag if active contacts < 2 (single-threaded) OR active < 3 OR all active contacts share one persona (under-threaded).
- Only contacts with is_former = false count as re-engagement targets; former contacts are excluded as targets.

Active-count arithmetic per flagged deal (rows − former − stale):

1) Deal-EC3025 (C-FDD0C7)
   Amount: not in data | Stage: not in data
   Active = 2 rows − 1 former (CT-F2C1AE, CPO) − 0 stale = 1  -> SINGLE-THREADED
   Personas present: champion (CT-047C54, 2026-09-02)
   Personas missing: economic buyer, HR admin, IT security, finance
   Most valuable add (stage unknown, so not stage-derived): economic buyer — no active EB on the deal and the only prior EB is former
   On-file unengaged fit: CT-6827DB (Chief People Officer, economic buyer)

2) Deal-92D97D (C-E23238)
   Amount: not in data | Stage: not in data
   Active = 2 rows − 0 former − 1 stale (CT-A902AE champion, 2026-06-01 < cutoff) = 1  -> SINGLE-THREADED
   Personas present: HR admin (CT-01F5B4, 2026-08-28)
   Personas missing: economic buyer, champion, IT security, finance
   Most valuable add (not stage-derived): champion — thread has no active champion and the prior one merely lapsed
   On-file unengaged fit: CT-A902AE (Head of Employee Experience, champion, stale since 2026-06-01, not former). For economic buyer: none on file.

3) Deal-50D386 (C-EB10E4)
   Amount: not in data | Stage: not in data
   Active = 2 rows − 0 − 0 = 2  -> UNDER-THREADED (< 3)
   Personas present: champion (2026-09-01), HR admin (2026-08-25)
   Personas missing: economic buyer, IT security, finance
   Most valuable add (not stage-derived): economic buyer — no EB in the thread
   On-file unengaged fit: CT-A1C4B3 (Chief People Officer, economic buyer)

4) Deal-D0D6B5 (C-32918E)
   Amount: not in data | Stage: not in data
   Active = 3 rows = 3, all champions  -> UNDER-THREADED (all one persona)
   Personas present: champion x3 (2026-09-02, 2026-08-19, 2026-08-07)
   Personas missing: economic buyer, HR admin, IT security, finance
   Most valuable add (not stage-derived): economic buyer — highest-authority persona entirely absent
   On-file unengaged fit: CT-1FA4DB (Chief People Officer, economic buyer)

5) Deal-5BFE3B (C-535D36)
   Amount: not in data | Stage: not in data
   Active = 2 rows = 2, all champions  -> UNDER-THREADED (< 3 AND all one persona)
   Personas present: champion x2 (2026-08-31, 2026-08-12)
   Personas missing: economic buyer, HR admin, IT security, finance
   Most valuable add (not stage-derived): economic buyer
   On-file unengaged fit: none on file (no C-535D36 rows in unengaged_contacts.csv; no stale EB anywhere in the deal file)

6) Deal-36C33F (C-077A0E)
   Amount: not in data | Stage: not in data
   Active = 3 rows − 2 former (CT-405B45 champion, CT-86B22F EB) − 0 = 1  -> SINGLE-THREADED
   Personas present: IT security (CT-4FE556, 2026-08-15)
   Personas missing: economic buyer, champion, HR admin, finance
   Most valuable add (not stage-derived): economic buyer — both EB and champion were marked former, so the deal has lost its buying committee; note the sole active contact is IT security, an unusual single-thread
   On-file unengaged fit: CT-1DB73E (Chief People Officer, economic buyer). For champion: none on file (former excluded).

7) Deal-885F45 (C-5E8EFB)
   Amount: not in data | Stage: not in data
   Active = 2 rows = 2  -> UNDER-THREADED (< 3)
   Personas present: economic buyer (2026-08-26), champion (2026-08-11)
   Personas missing: HR admin, IT security, finance
   Most valuable add (not stage-derived): cannot be ranked by stage; IT security is the only missing persona with an on-file contact
   On-file unengaged fit: CT-B3F25D (IT Security Lead, IT security). HR admin and finance: none on file.

8) Deal-FCBE5B (C-737030)
   Amount: not in data | Stage: not in data
   Active = 1 row = 1  -> SINGLE-THREADED
   Personas present: champion (CT-4A5317, 2026-08-29)
   Personas missing: economic buyer, HR admin, IT security, finance
   Most valuable add (not stage-derived): economic buyer — 1-contact champion-only thread
   On-file unengaged fit: none on file

9) Deal-5408B0 (C-2AE3AA)
   Amount: not in data | Stage: not in data
   Active = 2 rows = 2  -> UNDER-THREADED (< 3)
   Personas present: champion (2026-09-01), HR admin (2026-08-18)
   Personas missing: economic buyer, IT security, finance
   Most valuable add (not stage-derived): economic buyer
   On-file unengaged fit: CT-07FA76 (Chief People Officer, economic buyer)

10) Deal-C6D97A (C-5A8FC2)
    Amount: not in data | Stage: not in data
    Active = 3 rows = 3, all champions  -> UNDER-THREADED (all one persona)
    Personas present: champion x3 (2026-08-31, 2026-08-20, 2026-08-05)
    Personas missing: economic buyer, HR admin, IT security, finance
    Most valuable add (not stage-derived): economic buyer
    On-file unengaged fit: none on file

11) Deal-F9A08A (C-0D15DF)
    Amount: not in data | Stage: not in data
    Active = 2 rows − 0 former − 1 stale (CT-913581 EB, 2026-06-20 < cutoff) = 1  -> SINGLE-THREADED
    Personas present: champion (CT-931B10, 2026-09-03)
    Personas missing (active): economic buyer, HR admin, IT security, finance
    Most valuable add (not stage-derived): economic buyer — the deal has one, but it lapsed on 2026-06-20 and is not former
    On-file unengaged fit: CT-913581 (Chief People Officer, economic buyer, stale since 2026-06-20 — re-engage first); a second EB, CT-697541, is also on file

Not flagged (pass: >= 3 active, multi-persona): Deal-84DBA6 (3 active: champion/EB/IT security), Deal-4B0BEB (4 active: champion/EB/HR admin/finance), Deal-D348E1 (5 active: all five personas).

Totals: 14 deals — 4 single-threaded, 7 under-threaded, 3 pass; 11 flagged. Summary gap counts across the 11 flagged: economic buyer missing from active threads in 10 of 11 (all but Deal-885F45), finance in 11, IT security in 10.

Data needed to finish the ask as specified: an amount and stage column per deal_id — the stage-conditioned "most valuable persona to add" is not derivable from what was provided.
talk-track-from-transcriptsPASS 1.00
Plain-English question: From this rep's last 10 calls: how do they open, handle objections, and how often do they lock a next step?
Prompt — what the model saw
From this rep's last 10 calls: what they lead with in the first five minutes, how they handle the three most common objections, how often a concrete next step is agreed (rate), and every competitor a prospect raised. Quote at most one line per finding and only lines that appear in the transcripts. End with two coaching notes.
cost $0.0023 · 58s · in 3,391 / out 3,246 tok
✓ objection_count✓ next_step_rate✓ competitors✓ coaching✓ quote_used

✓ Correct Answer (ground truth)

Plain-English answer: 8 of 10 calls open with the same customer-win story; objection handling, next-step agreement rate, and every competitor raised, with quotes.
The answer in one line: n_transcripts = 10; objection_count = 3; rep = Alex Franklin (+3 more below)
n_transcripts10
objection_count3
allowed_new13 items — e.g. TT-001, TT-002, TT-003, TT-004
competitor_quotesAwardco=We're also in late talks with Awardco — their rewards catalog looks bigger than yours.; Kudos=How are you different from Kudos? Our CEO used them at her last company.
competitors_prospect_raisedAwardco, Kudos
next_stepagreed=7; total=10; rate=0.7
objectionsobjection=budget; count=4; quote=Honestly, budget is locked until next fiscal year — I can't add a new line item right now.; transcript_id=TT-001 | objection=timing; count=3; quote=This is good, but can we revisit it next quarter? Open enrollment starts in October and we're underwater.; transcript_id=TT-002 | objection=in-house; count=3; quote=We already do recognition with a spreadsheet and quarterly gift cards — why would we change?; transcript_id=TT-004
opener_themetheme=opens with a customer recognition win story (400-person retailer, 18% turnover reduction); count=8; total=10; quote=Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it.; transcript_id=TT-001
repAlex Franklin
rep_raised_competitors_excludedWorkhuman
graded checksobjection_count · next_step_rate · competitors · coaching · quote_used

Model's actual answer

show the model's full answer
CALL ANALYSIS — Alex Franklin, last 10 calls (TT-001 to TT-010, 2026-07-28 to 2026-09-03)

1. LEAD (first five minutes)
8 of 10 calls open with the identical retailer proof story; 1 opens with an agenda (Deal-403845), 1 with a pricing callback (Deal-1E2498). Count: 8 + 1 + 1 = 10.
Quote: "Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it."

2. THREE MOST COMMON OBJECTIONS AND HOW THEY'RE HANDLED
Counted from prospect lines: locked budget = 4 (Deal-D348E1, Deal-547B2B, Deal-60C2C2, Deal-84DBA6), revisit-next-quarter = 3 (Deal-5408B0, Deal-C61CF7, Deal-D9A12F), status-quo spreadsheet = 3 (Deal-403845, Deal-EDC141, Deal-1E2498).

- Locked budget (4/4 handled): reframed to turnover-savings funding with a $210k backfill figure. "Most teams fund this out of turnover savings — that retailer saved about $210k in avoided backfills, which is how their finance team signed off."
- Revisit next quarter (3/3 handled): countered with a scoped pilot to generate internal data. "What if we scope a 90-day pilot with one department so you have internal data before next quarter's planning?"
- Status-quo spreadsheet (3/3 handled): pivoted to automation + analytics vs manual effort. "Spreadsheets work until they scale — the difference is automation: milestones fire without HR lifting a finger, and you get analytics on who is being recognized."

Note: two lower-frequency objections got no handling — committee gates (2: Deal-403845, Deal-84DBA6) and "no urgency" (1: Deal-EDC141), met only with "I'll leave it with you." / "Fair enough."

3. CONCRETE NEXT STEP AGREED
Proposed as a working session in 7 of 10 calls, and accepted in all 7 (7/7 = 100% when proposed). Agreed rate: 7/10 = 70%. Missed in 3 of 10 (30%): Deal-403845, Deal-EDC141, Deal-84DBA6.
Quote: "Should we lock the next step — a working session with your team this week?"

4. COMPETITORS RAISED BY PROSPECTS
2 competitors, 2 calls (1 mention each):
- Awardco (Deal-547B2B): "We're also in late talks with Awardco — their rewards catalog looks bigger than yours."
- Kudos (Deal-EDC141): "How are you different from Kudos? Our CEO used them at her last company."
Not counted: Workhuman (Deal-C61CF7) was raised by the rep, not the prospect.

COACHING NOTES
1. The 3 zero-next-step calls (Deal-403845, Deal-EDC141, Deal-84DBA6) are exactly the calls where a committee or "no urgency" objection surfaced — that's where Alex goes passive ("Fair enough."). Treat a committee as a scheduling constraint, not a stop: agree a dated callback and ask who's on the committee, so the working session gets locked anyway.
2. The opener is delivered verbatim in 8 of 10 calls; the two adapted leads (security/pricing agenda, "You asked for straight pricing last time") show the range exists. Tailor the first five minutes to the account's actual context before the proof story, and prepare a distinct answer for catalog breadth (Awardco) — the current reply ("where we win is automation and the analytics") concedes the catalog point without evidence.
reporting-analytics 5 tests
q3-weighted-forecastPASS 1.00
Plain-English question: Build the Q3 2026 weighted forecast from the open pipeline, and show the math deal by deal.
Prompt — what the model saw
Below are the open deals in the default sales pipeline with a close date between 2026-07-01 and 2026-10-15, pulled from the CRM warehouse on 2026-09-05. Stages: DS1 (discovery) through DS5 (contract). Build the Q3 2026 forecast (quarter = 2026-07-01 to 2026-09-30). Weighting rule: 100 percent of COMMIT plus 35 percent of BEST_CASE. PIPELINE deals count zero. Only deals with a close date inside the quarter count. Report: COMMIT total, BEST_CASE total, weighted forecast, count of deals in each category inside the quarter, the deals you excluded for being outside the quarter (count and total amount), and the top 5 BEST_CASE deals by amount inside the quarter. Show arithmetic. Then, in a final '## Data quality' section of no more than four sentences, name the data quality problems in this extract that would make the forecast unreliable if it were run unattended.
cost $0.0093 · 113s · in 30,671 / out 9,237 tok
✓ commit_total_44729✓ best_case_total_203565✓ weighted_115976.75✓ commit_count_7✓ best_case_count_24✓ excluded_count_32✓ excluded_amount_227575✓ top5_aliases✓ flags_blank_owner✓ flags_missing_why_buys

✓ Correct Answer (ground truth)

Plain-English answer: A stage-weighted Q3 forecast number computed from the open pipeline, with per-stage math that reconciles to the deal data.
The answer in one line: excluded_amount = 227575; best_case_total = 203565; weighted_forecast = 115,976.75; commit_total = 44729; excluded_out_of_quarter = 32; best_case_count = 24; flags = blank owner column + empty why_buys fields must be surfaced (+2 more below)
excluded_amount227,575
best_case_count24
best_case_total203,565
commit_count7
commit_total44,729
excluded_out_of_quarter32
weighted_forecast115,976.75
flagsblank owner column + empty why_buys fields must be surfaced
top5_best_caseDeal-2D7423, Deal-25F752, Deal-E53952, Deal-5EED42, Deal-FA32A0
graded checkscommit_total_44729 · best_case_total_203565 · weighted_115976.75 · commit_count_7 · best_case_count_24 · excluded_count_32 · excluded_amount_227575 · top5_aliases · flags_blank_owner · flags_missing_why_buys

Model's actual answer

show the model's full answer
Q3 2026 FORECAST (quarter = 2026-07-01 to 2026-09-30; extract of 86 open deals pulled 2026-09-05)

COMMIT (in-quarter) — 7 deals
  Deal-547B2B 11,200 + Deal-B7EBD1 9,000 + Deal-403845 9,000 + Deal-A2B47C 6,360 + Deal-2465CE 5,400 + Deal-A5E80A 2,520 + Deal-499BF6 1,249
  = 44,729

BEST_CASE (in-quarter) — 24 deals
  Total = 203,565

PIPELINE (in-quarter) — 23 deals, 201,637.40 — counts 0.00 by rule.

WEIGHTED FORECAST
  = 100% COMMIT + 35% BEST_CASE
  = 44,729 + 0.35 × 203,565
  = 44,729 + 71,247.75
  = 115,976.75

Deal counts inside the quarter: COMMIT 7, BEST_CASE 24, PIPELINE 23 (zero-weighted), total in-quarter 54.

EXCLUDED AS OUTSIDE THE QUARTER — 32 deals, 227,575.00 total
  All have close dates 2026-10-01 to 2026-10-15 (none before 2026-07-01).
  By category: PIPELINE 22 deals / 185,565; BEST_CASE 9 deals / 28,240 (Deal-C61CF7 5,400; Deal-48B656 5,160; Deal-901332 3,600; Deal-47AE31 3,600; Deal-15D24F 3,600; Deal-ED725A 2,400; Deal-8AD4A5 1,800; Deal-5FDCE4 1,600; Deal-F5A622 1,080); COMMIT 1 deal / 13,770 (Deal-D348E1, DS5, closes 2026-10-15).

TOP 5 BEST_CASE DEALS IN-QUARTER
  1. Deal-2D7423 — 38,935 (DS3, 2026-09-30)
  2. Deal-25F752 — 24,000 (DS4, 2026-09-25)
  3. Deal-E53952 — 19,656 (DS4, 2026-09-30)
  4. Deal-5EED42 — 16,250 (DS3, 2026-09-30)
  5. Deal-FA32A0 — 11,116 (DS3, 2026-09-25)

## Data quality
Owner is blank on 85 of 86 rows and why_buys is 0 characters on 71 of 86, so there is no accountability or win-rationale to sanity-check category assignments. Forecast_category contradicts stage in several places — Deal-A5E80A is COMMIT at DS1 and Deal-499BF6 is COMMIT at DS2, while DS4's Deal-6787C2 is only PIPELINE — suggesting stale or miscategorized fields that directly move the weighted number. Close dates cluster suspiciously: 17 deals sit on 2026-09-30 (quarter boundary) and 11 more on 2026-10-15 (just outside), so small date slips shift large sums across the quarter line, including Deal-D348E1's 13,770 COMMIT. Concentration and hygiene compound the risk: Deal-2D7423 alone is 19% of the in-quarter BEST_CASE total and ~12% of the whole weighted forecast, one amount (2,480.4 for Deal-42326B) is non-integer, and the extract contains no currency field or created/updated timestamps to detect duplicates or aged records.
aha-moment-2x20.83
Plain-English question: Do companies that hit 5+ givers and one redemption in month one actually retain better at 24 months?
Prompt — what the model saw
Test the activation hypothesis on the company cohort below: companies with 5+ unique givers (m1_users >= 5) AND 1+ successful redemption (m1_redemptions >= 1) in their first calendar month retain better at 24 months than those with one signal or neither. Every company in this extract is already 25+ months old, so current_status = 'active' means retained at 24 months. Report the full 2x2 (cohort size and 24-month retention rate for: both signals, givers-only, redemption-only, neither), state how many companies were excluded from the denominator and why (if none, say so), name the single signal with the largest retention lift, and state what this does and does not prove.
cost $0.0127 · 186s · in 26,125 / out 15,462 tok
✓ both_66✓ givers_47✓ redemp_31✓ neither_40✗ excluded_0✓ largest_lift_givers

✓ Correct Answer (ground truth)

Plain-English answer: The four cohort cells are graded against computed 24-month retention; the numbers come from the scorer's cohort table, not opinion.
The answer in one line: both_5plus_givers_and_redemption = 66.0% (31 of 47)
both_5plus_givers_and_redemption66.0% (31 of 47)
exclusionsnone — claim bounded as correlation, not causation
givers_only46.9% (23 of 49)
largest_single_signal_liftunique givers
neither40.0% (38 of 95)
redemption_only31.0% (9 of 29)
graded checksboth_66 · givers_47 · redemp_31 · neither_40 · excluded_0 · largest_lift_givers

Model's actual answer

show the model's full answer
Verified against the full extract (220 rows parsed; the file was read from /Users/amaniphipps/plg_company_cohort_2023.csv, a verbatim copy of the attachment).

ACTIVATION HYPOTHESIS TEST — 2023 cohort, 24-month retention
Definition: retained = current_status == 'active'. Signals: m1_users >= 5 (givers), m1_redemptions >= 1.

FULL 2x2
  Cell                         n    active   retention
  Both signals                 47      31     31/47 = 66.0%
  Givers-only                  49      23     23/49 = 46.9%
  Redemption-only              29       9      9/29 = 31.0%
  Neither                      95      38     38/95 = 40.0%
  Total                       220     101    101/220 = 45.9%

EXCLUSIONS FROM DENOMINATOR
  Zero companies excluded. All 220 rows have valid m1_users, m1_redemptions, and current_status, and every row lands in exactly one cell (47+49+29+95 = 220; 31+23+9+38 = 101 active, matching the status count). Judgment call to flag: the 3 'non_renewing' companies (C-0B2078FB, C-0A96134F, C-0BEAF685) are counted as NOT retained, since the given rule defines retention strictly as active. Two of them sit in the redemption-only cell, so that cell's 31.0% is slightly depressed by that choice.

VERDICT ON THE HYPOTHESIS
  Partly supported, not cleanly. The both-signals cell is the best performer (66.0% vs 40.0% for neither, +26.0 pts). But the ordering is NOT monotonic: givers-only (46.9%) beats neither, while redemption-only (31.0%) is 9.0 pts BELOW neither. Having a redemption without 5+ users is associated with worse-than-baseline retention, not better.

LARGEST SINGLE-SIGNAL LIFT: unique givers (m1_users >= 5)
  Marginal test (signal present vs absent, other signal ignored):
    Givers present:     (31+23)/(47+49) = 54/96 = 56.3%
    Givers absent:      (9+38)/(29+95)  = 47/124 = 37.9%   → lift = +18.3 pts
    Redemption present: (31+9)/(47+29)  = 40/76 = 52.6%
    Redemption absent:  (23+38)/(49+95) = 61/144 = 42.4%   → lift = +10.3 pts
  Isolated-cell test (vs neither): givers-only +6.9 pts; redemption-only −9.0 pts.
  Under either reading, givers is the stronger signal.

WHAT THIS DOES PROVE
  - A real descriptive association in this cohort: early multi-user adoption co-occurs with much higher 24-month survival. The both-vs-neither gap (66.0% vs 40.0%) is beyond noise at these sizes (z ≈ 3.0, p < 0.01).
  - The signals interact positively: both (66.0%) exceeds either alone, so redemption adds value mainly on top of a >=5-user footprint.

WHAT IT DOES NOT PROVE
  - Causality. m1_users is a proxy for company size and buying intent — larger logo deals (e.g. C-0D0B047C, C-22170CA1) both activate fast and were always likelier to stick. No experiment, no control for tier/industry/country/ARR, so the "activation → retention" link could be selection, not treatment.
  - That the one-signal cells behave as the hypothesis predicts — redemption-only actively contradicts it (small accounts that redeemed once churn at the worst rate; see C-F0A562FF: 0 users, 10 redemptions, cancelled, and C-0F58EBD7: 1 user, 42 redemptions, cancelled). Redemption without adoption breadth is not a save signal.
  - Precision of the single-signal gaps: givers-only vs neither (+6.9 pts, n=49) and redemption-only vs neither (−9.0 pts, n=29) are well within sampling noise on their own.
  - Anything about the 2023 cohort's thresholds generalizing — 5 users is an arbitrary cut applied to a 220-company extract; the effect could be a size effect wearing an activation costume.
arr-rollforward-reconciliationPASS 1.00
Plain-English question: Does CRM ARR match billing ARR — and if not, exactly where is the variance?
Prompt — what the model saw
Reconcile ARR between CRM company records and billing subscriptions as of 2026-09-05. Billing ARR is MRR x 12. Report both totals, the variance, and a decomposition into buckets (status mismatch, rounding, missing records, other) that sums exactly to the variance. List each mismatched account with a suggested owner. Business rule: any subscription with a term other than 12 months must have cf_agreement_end_date populated; list violations.
cost $0.0052 · 128s · in 5,664 / out 7,914 tok
✓ billing_total✓ crm_total✓ variance✓ buckets_sum✓ violations_listed

✓ Correct Answer (ground truth)

Plain-English answer: Billing $604,739.28 vs CRM $603,581.76: a $1,157.52 variance decomposed into status-mismatch, rounding, and missing-record buckets.
The answer in one line: n_companies = 40; n_mismatched_accounts = 7; as_of = 2026-09-05 (+4 more below)
n_companies40
n_mismatched_accounts7
as_of2026-09-05
billing_arr_total604739.28
bucket_sum_equals_variance1
buckets{"status_mismatch": {"amount": "-13158.48", "accounts": [{"company_alias": "C-0C8323BF", "billing_arr": "0.00", "hubspot_arr": "4905.24", "diff": "-4905.24", "suggested_owner": "Dana Mercer"}, {"compa
crm_arr_total603581.76
ruleterm_months != 12 requires cf_agreement_end_date
variance1157.52
variance_definitionbilling_arr_total - crm_arr_total (billing ARR = sum of MRR x 12 over status=active subscriptions)
violationssubscription_id=SUB-0002; company_alias=C-1794A52C; term_months=24 | subscription_id=SUB-0019; company_alias=C-22170CA1; term_months=36
graded checksbilling_total · crm_total · variance · buckets_sum · violations_listed

Model's actual answer

show the model's full answer
Reconciliation complete. All figures computed with exact decimal arithmetic; no external data used.

## 1. Totals

| Source | ARR |
|---|---|
| CRM (company_arr.csv, 39 records) | **$603,581.76** |
| Billing, active subscriptions only (37 subs × MRR×12) | **$604,739.28** |
| Variance (CRM − Billing active) | **−$1,157.52** (CRM is under) |

Note: billing all-in including the 2 cancelled subs = $617,897.76. I use the active-only baseline below; the cancelled subs are broken out as their own bucket.

## 2. Bucket decomposition (sums exactly to −$1,157.52)

Convention: each bucket = CRM minus Billing-active contribution.

**Bucket A — Status mismatch: +$13,158.48**
CRM still carries ARR for subscriptions that are `cancelled` in billing:
- C-0C8323BF: CRM $4,905.24 vs billing ARR $0 (cancelled, MRR 408.77×12 = 4,905.24) → +4,905.24
- C-0DC4FB8C: CRM $8,253.24 vs billing ARR $0 (cancelled, MRR 687.77×12 = 8,253.24) → +8,253.24

**Bucket B — Missing records: −$11,952.00**
- C-21629AA4 (SUB-0004): active in billing, MRR 2,370.77 → ARR 2,370.77×12 = $28,449.24, but NO company record in CRM → −28,449.24
- C-0D5BBE3A: in CRM with $16,497.24 but no subscription in billing → +16,497.24
- Net: −28,449.24 + 16,497.24 = −11,952.00

**Bucket C — Rounding: +$36.00**
- C-0D66DF9E (SUB-0005): billing 1,932.00×12 = $23,184.00 vs CRM $23,200.00 → +16.00 (CRM figure is a rounded ARR: 23,200/12 = 1,933.33)
- C-14D70CE0 (SUB-0008): billing 1,515.00×12 = $18,180.00 vs CRM $18,200.00 → +20.00 (18,200/12 = 1,516.67)

**Bucket D — Other: −$2,400.00**
- C-0F7269D7 (SUB-0006): billing 2,233.00×12 = $26,796.00 vs CRM $24,396.00 → −2,400.00. CRM implies MRR of 24,396/12 = $2,033.00 — a $200/mo gap, consistent with a stale MRR in CRM rather than rounding.

**Check: +13,158.48 − 11,952.00 + 36.00 − 2,400.00 = −$1,157.52 ✓**

## 3. Mismatched accounts with suggested owner

The provided data contains no owner/assigned-to field, so I cannot name individuals — suggesting owners by function (bucket accountability):

| Account | Issue | Δ (CRM − billing ARR) | Suggested owner |
|---|---|---|---|
| C-0C8323BF | Status: cancelled in billing, ARR live in CRM | +4,905.24 | CS/RevOps — remove closed-won ARR on churn |
| C-0DC4FB8C | Status: cancelled in billing, ARR live in CRM | +8,253.24 | CS/RevOps |
| C-21629AA4 | Missing: active sub, no CRM company record | −28,449.24 | CRM data ops — create/link company record |
| C-0D5BBE3A | Missing: CRM ARR, no subscription | +16,497.24 | RevOps — verify deal source or zero the field |
| C-0F7269D7 | Other: CRM MRR likely stale ($2,033 vs $2,233) | −2,400.00 | Sales ops — sync current MRR to CRM |
| C-0D66DF9E | Rounding (manual rounded ARR) | +16.00 | RevOps — set CRM ARR = MRR×12 |
| C-14D70CE0 | Rounding (manual rounded ARR) | +20.00 | RevOps — same |

## 4. Business-rule violations — non-12-month term without cf_agreement_end_date

| Subscription | Company | Term | End date | Verdict |
|---|---|---|---|---|
| SUB-0002 | C-1794A52C | 24 mo | (empty) | **VIOLATION** |
| SUB-0019 | C-22170CA1 | 36 mo | (empty) | **VIOLATION** |
| SUB-000C | C-0DB48281 | 24 mo | 2027-11-30 | compliant |
| SUB-001A | C-0FC4DBB8 | 36 mo | 2027-11-30 | compliant |

2 violations out of 4 non-12-month subscriptions. (Observation, outside the rule as written: both populated end dates are 2027-11-30, which is <24 months for SUB-001A's 36-month term and matches SUB-000C's 24-month window only loosely — worth a date-accuracy spot check.)

All 32 remaining matched active accounts reconcile to $0.00 difference.
kvm-month-over-monthPASS 1.00
Plain-English question: How did the four core KVMs move in August versus July?
Prompt — what the model saw
Report the four core KVMs for 2026-08 against 2026-07: giving rate, redemptions per user, 1:1 meetings engagement, pulse check engagement. For each: value, prior value, absolute and relative change, direction. Name the KVM with the largest relative move and the segment (plan_tier or size_band) driving it if the data supports that; otherwise say the data does not. Under 200 words plus one table.
cost $0.0081 · 167s · in 7,689 / out 13,278 tok
✓ four_kvms✓ biggest_mover_named✓ direction_correct✓ segment_named✓ word_cap

✓ Correct Answer (ground truth)

Plain-English answer: Giving rate 60.27% (up 0.07%), redemptions per user 1.7302 (up 0.01%), with 1:1 meetings and pulse engagement each up fractionally — direction reported for all four.
The answer in one line: n_company_rows_per_month = 30; aggregation = simple mean of company-level values per month
n_company_rows_per_month30
aggregationsimple mean of company-level values per month
biggest_moverpulse_engagement
driving_segmentsize_band=enterprise; 2026-07=0.55; 2026-08=0.2743; rel_change_pct=-50.13
kvms{"giving_rate": {"2026-07": 0.6023, "2026-08": 0.6027, "abs_change": 0.0004, "rel_change_pct": 0.07, "direction": "up"}, "redemptions_per_user": {"2026-07": 1.73, "2026-08": 1.7302, "abs_change": 0.00
segment_detail{"smb": {"2026-07": 0.6588, "2026-08": 0.6573, "rel_change_pct": -0.23}, "mid_market": {"2026-07": 0.593, "2026-08": 0.5942, "rel_change_pct": 0.2}, "enterprise": {"2026-07": 0.55, "2026-08": 0.2743,
graded checksfour_kvms · biggest_mover_named · direction_correct · segment_named · word_cap

Model's actual answer

show the model's full answer
All 30 companies are tier_three, so segment attribution uses size_band only. Values are unweighted means across the 30 company aliases (10 per band).

| KVM | 2026-07 | 2026-08 | Abs Δ | Rel Δ | Direction |
|---|---|---|---|---|---|
| Giving rate | 0.6023 | 0.6027 | +0.0004 | +0.07% | up |
| Redemptions per user | 1.7300 | 1.7302 | +0.0002 | +0.01% | flat/up |
| 1:1 meetings engagement | 0.4469 | 0.4472 | +0.0003 | +0.06% | up |
| Pulse check engagement | 0.6006 | 0.5086 | −0.0920 | −15.31% | down |

Arithmetic: e.g. pulse July mean 0.6006, August mean 0.5086; abs = 0.5086 − 0.6006 = −0.0920; rel = −0.0920 / 0.6006 = −15.31%.

Largest relative move: pulse check engagement (−15.31%). The data supports the segment: enterprise size_band. All 10 enterprise companies (C-0B2895EF, C-0B2213A9, C-0D6CC8E3, C-0D0B047C, C-0D3278C7, C-0FCCD2DF, C-0F6C0F34, C-8C2E8F00, C-0B827671, C-0BA71F12) fell from a 0.5331–0.5661 range in July to 0.2619–0.2864 in August — roughly −50% each. Enterprise band mean: 0.5500 → 0.2743 (−50.1%). SMB (0.6582 → 0.6576) and mid_market (0.5930 → 0.5948, +0.3%) were flat, so the entire portfolio decline is enterprise-driven. The uniform, ~50%-cut pattern across every enterprise account suggests a measurement/onboarding change rather than organic behavior, but the data provided does not confirm a cause.
rewards-ytd-sectionPASS 1.00
Plain-English question: Produce the redemption section of the weekly rewards report through August.
Prompt — what the model saw
Produce the redemption section of the weekly rewards report through the last completed month (2026-08): redemption count, spend, unique redeemers, redemptions per redeemer, provider mix as percent of spend, top 5 countries by redemptions. Say which month is the last completed month. Provider shares must sum to 100.
cost $0.0382 · 550s · in 75,002 / out 49,674 tok
✓ count✓ spend✓ unique_redeemers✓ month_named✓ provider_mix✓ top_country

✓ Correct Answer (ground truth)

Plain-English answer: 378 redemptions, $27,846 spend, 235 unique redeemers at 1.61 each, custom rewards at 39% of spend, US leading all countries.
The answer in one line: redemption_count = 378; unique_redeemers = 235; redemptions_per_redeemer = 1.61; total_spend_usd = 27846.00
redemption_count378
redemptions_per_redeemer1.61
unique_redeemers235
provider_mix_pct_of_spendTangoCard=11.6; Tremendous=30.54; custom=39.05; Snappy=18.81
top5_countries_by_countcountry=US; count=244 | country=CA; count=24 | country=AU; count=21 | country=GB; count=17 | country=NL; count=17
total_spend_usd27846.00
graded checkscount · spend · unique_redeemers · month_named · provider_mix · top_country

Model's actual answer

show the model's full answer
Redemption section — through the last completed month. Today is 2026-09-25, so the last completed month is **August 2026 (2026-08)**. Window = 2026-01-01 through 2026-08-31. The file's latest record is RD-B7BEED at 2026-08-31T09:10:00, so all 378 rows in redemptions_ytd.csv fall in scope; none excluded.

SCOPE
- Rows parsed: 378 unique redemption_ids, all dated 2026-01 through 2026-08.
- Monthly counts (sum = 378): Jan 45, Feb 36, Mar 50, Apr 52, May 43, Jun 46, Jul 58, Aug 48.
  45+36+50+52+43+46+58+48 = 378 ✓
- Monthly spend (sum = $27,846.00): Jan $3,932 + Feb $2,731 + Mar $2,918 + Apr $3,876 + May $2,825 + Jun $2,773 + Jul $5,255 + Aug $3,536 = $27,846 ✓

HEADLINE METRICS
- Redemption count: **378**
- Spend: **$27,846.00** (amount_usd summed directly; no state/validity filter exists in this file to apply)
- Unique redeemers: **235** distinct user_key values, deduplicated across the whole window (not the sum of monthly uniques)
- Redemptions per redeemer: 378 ÷ 235 = **1.61** (1.6085)

PROVIDER MIX (% of spend; shares at 2dp sum to exactly 100.00)
| Provider | Spend | Count | % of spend |
|---|---|---|---|
| custom | $10,873.00 | 37 | 10873/27846 = 39.05% |
| Tremendous | $8,505.00 | 192 | 8505/27846 = 30.54% |
| Snappy | $5,238.00 | 59 | 5238/27846 = 18.81% |
| TangoCard | $3,230.00 | 90 | 3230/27846 = 11.60% |
| Total | $27,846.00 | 378 | 100.00% |

Check: 10,873 + 8,505 + 5,238 + 3,230 = 27,846 ✓; 39.05 + 30.54 + 18.81 + 11.60 = 100.00 ✓

TOP 5 COUNTRIES BY REDEMPTIONS
1. US — 244
2. CA — 24
3. AU — 21
4. GB — 17 (tie)
4. NL — 17 (tie)
Top-5 subtotal: 244+24+21+17+17 = 323 of 378 (85.4%). GB and NL tie at 17 for 4th/5th; next country SG at 12 is clear of the cut, so the top-5 set is unambiguous.

Notes on data limits: provider roster taken exactly as labeled in the file (TangoCard, Tremendous, custom, Snappy — "custom" lowercase as given); no redemption-record normalization applied since the file already carries provider. Amounts are USD in the source column amount_usd.
customer-success 5 tests
churn-save-eligibilityPASS 1.00
Plain-English question: Which at-risk accounts actually qualify for a churn-save offer, and which play fits each?
Prompt — what the model saw
Which at-risk accounts qualify for a churn-save offer under the documented eligibility rules, what amount is at stake per account and in total, and which play fits each (usage revival, executive touch, commercial concession)? Cite the signal that justifies each play. List accounts that look at risk but do not qualify and why.
cost $0.0060 · 153s · in 4,441 / out 10,009 tok
✓ eligible_set✓ total_at_stake✓ plays_cited✓ noneligible_named✓ rules_applied

✓ Correct Answer (ground truth)

Plain-English answer: Only accounts meeting all three rules qualify — e.g. C-0F6C0F34 at $49,707 (executive touch) and C-0B827671 at $25,365 (usage revival) — with amount and play per account.
The answer in one line: n_accounts = 30; snapshot_date = 2026-09-05
n_accounts30
eligibleaccount_alias=C-0F6C0F34; amount_at_stake=49707.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0B827671; amount_at_stake=25365.00; play=usage revival; justifying_signal=usage_trend_3m=declining over the last 3 months | account_alias=C-0B360C78; amount_at_stake=35748.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 75% - risk is commercial, not adoption | account_alias=C-0B0F1BAB; amount_at_stake=5494.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0CA21961; amount_at_stake=16829.00; play=usage revival; justifying_signal=seat utilization 26% is below 50% | account_alias=C-0E9C27D1; amount_at_stake=41235.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 85% - risk is commercial, not adoption … (+2 more)
non_eligible_at_riskaccount_alias=C-0BC71BDD; health_score=55 | account_alias=C-0BA71F12; health_score=52 | account_alias=C-0F6694C3; health_score=43 | account_alias=C-0BE96399; health_score=54 | account_alias=C-0F876796; health_score=47 | account_alias=C-0FCCD2DF; health_score=43 … (+1 more)
ruleshealth_score < 60, churn_save_eligible_amount > 0, renewal within 120 days of snapshot
snapshot_date2026-09-05
total_amount_at_stake224601.00
graded checkseligible_set · total_at_stake · plays_cited · noneligible_named · rules_applied

Model's actual answer

show the model's full answer
Eligibility pass (snapshot 2026-09-05; window = renewal within 120 days → on or before 2027-01-03; R1 health<60 AND R2 churn_save_eligible_amount>0 AND R3 renewal in window)

QUALIFY — 8 accounts (all pass R1+R2+R3)

| Account | Health | Renewal (days out) | ARR | Save amount |
|---|---|---|---|---|
| C-0F6C0F34 | 51 | 2026-10-03 (28d) | 86,741.00 | 49,707.00 |
| C-0B827671 | 56 | 2026-11-14 (70d) | 72,088.00 | 25,365.00 |
| C-0B360C78 | 57 | 2026-10-28 (53d) | 60,427.00 | 35,748.00 |
| C-0B0F1BAB | 38 | 2026-09-23 (18d) | 15,391.00 | 5,494.00 |
| C-0CA21961 | 58 | 2026-12-28 (114d) | 31,501.00 | 16,829.00 |
| C-0E9C27D1 | 39 | 2026-09-24 (19d) | 75,093.00 | 41,235.00 |
| C-0CEF69FD | 53 | 2026-11-21 (77d) | 79,324.00 | 32,621.00 |
| C-0D3278C7 | 54 | 2026-11-12 (68d) | 33,815.00 | 17,602.00 |

Arithmetic — total at stake:
- Save pool: 49,707 + 25,365 + 35,748 + 5,494 + 16,829 + 41,235 + 32,621 + 17,602 = 224,601.00
- ARR exposed: 86,741 + 72,088 + 60,427 + 15,391 + 31,501 + 75,093 + 79,324 + 33,815 = 454,380.00

Play assignment per account. Caveat: the provided data contains no documented signal→play mapping rules; these assignments are analyst judgment, with the justifying signal cited from the files.

- C-0F6C0F34 — Executive touch. Signal: champion_active=false with usage healthy (growing, 308/395 seats ≈ 78%); renewal in 28 days needs a relationship above the missing champion.
- C-0B827671 — Usage revival. Signal: usage_trend_3m=declining, 113/202 seats ≈ 56% utilized, with an active champion to drive adoption.
- C-0B360C78 — Commercial concession. Signal: usage_trend_3m=growing at 246/327 ≈ 75% and champion_active=true — no usage or relationship gap to fix; the health-57 risk is commercial, and 35,748 of 60,427 ARR (≈59%) is at stake at renewal in 53 days.
- C-0B0F1BAB — Executive touch. Signal: champion_active=false, lowest health in the file (38), renewal in 18 days — the shortest runway; no internal advocate left to save it.
- C-0CA21961 — Usage revival. Signal: worst utilization among qualifiers, 84/325 seats ≈ 26%, trend flat; champion is active to run the adoption push.
- C-0E9C27D1 — Commercial concession. Signal: 134/157 seats ≈ 85% utilized, trend flat, champion active, health 39 — the account is used hard yet still at risk, and 41,235 of 75,093 (≈55%) is eligible; renewal in 19 days. Usage/relationship levers are exhausted; price/terms is what's left.
- C-0CEF69FD — Executive touch. Signal: champion_active=false with healthy usage (growing, 97/136 ≈ 71%).
- C-0D3278C7 — Usage revival. Signal: usage_trend_3m=declining and 126/380 seats ≈ 33% utilized — heaviest seat surplus of the qualifiers; champion active.

AT RISK (health<60) BUT NOT ELIGIBLE — 7 accounts

| Account | Health | Save amount | Renewal (days out) | Failed rule(s) | Why |
|---|---|---|---|---|---|
| C-0BC71BDD | 55 | 0.00 | 2026-10-27 (52d) | R2 | eligible amount is 0 |
| C-0BE96399 | 54 | 0.00 | 2026-10-29 (54d) | R2 | eligible amount is 0 |
| C-10A56B0F | 54 | 0.00 | 2026-12-12 (98d) | R2 | eligible amount is 0 |
| C-0F6694C3 | 43 | 0.00 | 2027-03-21 (197d) | R2, R3 | zero amount and renewal past 2027-01-03 |
| C-0FCCD2DF | 43 | 0.00 | 2027-04-23 (230d) | R2, R3 | zero amount and renewal past 2027-01-03 |
| C-0BA71F12 | 52 | 6,824.00 | 2027-04-11 (218d) | R3 | renewal outside 120-day window; enters it 2026-12-12 |
| C-0F876796 | 47 | 19,958.00 | 2027-02-06 (154d) | R3 | renewal outside window; enters it 2026-10-09 |

Note on usage signals for the non-qualifiers: C-0BE96399 (43/154 ≈ 28%, declining), C-0BC71BDD (59/197 ≈ 30%), C-0F876796 (22/95 ≈ 23%, declining), and C-0BA71F12 (23/98 ≈ 23%, declining) are textbook usage-revival profiles that carry no churn-save budget or window under R2/R3 — they need a budget exception, not an offer.

The remaining 15 accounts (health 62–88, e.g. C-8C2E8F00, C-14264ABD, C-0EE1203E) fail R1 and are not at risk; C-0D035262 renews in-window (2026-10-20) with healthy usage (245/262 ≈ 94%) but at $0 eligible.

Bottom line: 8 qualifying accounts, 224,601.00 in churn-save eligible amount against 454,380.00 ARR; two renewals (C-0B0F1BAB 18d, C-0E9C27D1 19d) are inside three weeks and should move first.
expansion-kitPASS 1.00
Plain-English question: Build the expansion kit for the customer who just replied to our outreach.
Prompt — what the model saw
A customer replied to outreach (reply included). Build the expansion kit: seat coverage (licensed over headcount), two lines on usage health, headroom in seats and ARR at the current per-seat rate, who replied and whether they can buy, the right buyer from our contacts, and a reply email under 150 words that moves toward the right person without being pushy and cites one real usage fact.
cost $0.0022 · 64s · in 662 / out 3,795 tok
✓ coverage_pct✓ headroom_arr✓ right_buyer✓ replier_not_buyer✓ usage_fact✓ email_under_150

✓ Correct Answer (ground truth)

Plain-English answer: 37.5% seat coverage, 250 headroom seats worth $15,000 ARR at $60 a seat, usage up 43.2% in six months — and the replier cannot buy; Dana R. (VP People) is the right buyer.
The answer in one line: headroom_seats = 250; coverage_pct = 37.5; company_alias = C-0DDFC9A7 (+3 more below)
headroom_seats250
coverage_pct37.5
company_aliasC-0DDFC9A7
headroom_arr15000.00
per_seat_annual_rate60.00
replierMaria S.
replier_can_buy0
right_buyername=Dana R.; title=VP People
usage_factsactive_users_2026_08=126; active_users_2026_03=88; growth_pct_6m=43.2; licensed_seat_utilization_pct=84
usage_health_linesMonthly active users grew from 88 (2026-03) to 126 (2026-08), +43.2% over 6 months., Last month 126 of 150 licensed seats were active (84.0% utilization), approaching the seat cap.
graded checkscoverage_pct · headroom_arr · right_buyer · replier_not_buyer · usage_fact · email_under_150

Model's actual answer

show the model's full answer
EXPANSION KIT — C-0DDFC9A7

1. SEAT COVERAGE
- Licensed 150 / headcount 400 = 150 ÷ 400 = 37.5% of employees covered by license.
- August usage 126 / 400 headcount = 31.5% of employees active.

2. USAGE HEALTH (two lines)
- Steady month-over-month growth with no dip: Mar 88 → Apr 95 → May 102 → Jun 110 → Jul 118 → Aug 126; +38 users over 5 steps = +43.2% (38 ÷ 88) and an average of +7.6 users/month.
- Consumption is tight: 126 ÷ 150 = 84.0% of licensed seats in use in August; 24 free seats ÷ 7.6/month ≈ 3.2 months of runway before the license is full.

3. HEADROOM AT CURRENT PER-SEAT RATE
- Derived rate: $9,000 ARR ÷ 150 seats = $60/seat/year. (Assumes ARR = seats × uniform rate; no price book given.)
- Seat headroom to 100% of headcount: 400 − 150 = 250 seats.
- ARR headroom: 250 × $60 = $15,000 incremental; full-coverage target = $24,000.

4. WHO REPLIED / CAN SHE BUY
- Maria S., People Operations Coordinator (last engaged 2026-09-02 — freshest of the three).
- Buying authority: not stated in the data. A coordinator title is a champion/influencer, not an economic buyer; treat her as the door, not the decision. No budget/approval data provided.

5. RIGHT BUYER FROM OUR CONTACTS
- Dana R., VP People — highest people-function authority on file, which matches a seat-expansion ask. Caveat: last engaged 2026-05-18, ~4 months stale; Maria's reply is the warm path in.
- Sam K. (Office Manager, last engaged 2025-11-03) — lowest authority and stalest; not the buyer.

6. REPLY DRAFT (115 words)

Subject: Your usage is about to hit the seat limit

Hi Maria,

Thanks for getting back to me.

One thing stood out in your account: monthly active users climbed from 88 in March to 126 in August — about a 43% increase, and roughly 84% of your 150 licensed seats. At that pace you'd fill the license in about three months.

No urgency from our side, but I'd rather you not get caught short. Is Dana R. the right person for a conversation like this, or does that budget sit elsewhere? Happy to send her a short summary instead of a meeting if that's easier.

Either way, glad to see the team using it this much.

Best,
[Name]

MISSING DATA (flagged, not assumed): Maria's actual budget authority; contract end date, renewal terms, and whether $60/seat is list or contract rate; usage after August 2026; headcount composition (how many of the 400 are in-scope for recognition licenses).
onboarding-statusPASS 1.00
Plain-English question: What is done and what is missing in this customer's onboarding before the mid-onboarding call?
Prompt — what the model saw
Prep the CSM for the mid-onboarding call: what is complete (integrations connected, allowance set, admins added, first recognition given, first redemption), what is not, early engagement signals, and the three things to cover on the call. Mark nothing complete without a data field that shows it.
cost $0.0019 · 54s · in 741 / out 3,197 tok
✓ complete_items✓ incomplete_items✓ three_items✓ grounded

✓ Correct Answer (ground truth)

Plain-English answer: 25 days live: Slack connected, allowance set, admins added — but the HRIS integration is not connected, with early engagement signals and three call priorities.
The answer in one line: days_live = 25; company_alias = C-0D284E42
days_live25
call_itemsConnect the HRIS integration (integration_hris is empty) to automate user provisioning and offboarding., Get to first redemption (first_redemption_at is empty): walk admins through the rewards catalog and confirm allowance funding., Build the giving habit: daily active givers averaged 4.3 in week 1 vs 13.0 last week - agree on manager nudges and a recognition cadence.
checklist{"slack_integration_connected": {"complete": true, "field": "integration_slack", "value": "2026-08-12"}, "hris_integration_connected": {"complete": false, "field": "integration_hris", "value": ""}, "a
company_aliasC-0D284E42
engagement_signalsavg_active_givers_week1=4.3; avg_active_givers_last_week=13; latest_active_givers=15; days_with_activity=25
graded checkscomplete_items · incomplete_items · three_items · grounded

Model's actual answer

show the model's full answer
CSM PREP — ACCOUNT C-0D284E42
Data window: signup 2026-08-11 through 2026-09-04 (last usage row). No data after 09-04 — it is 2026-09-25 today, so the usage feed is 20+ days stale. Flag noted below.

WHAT IS COMPLETE (each backed by a populated field)
- Slack integration: integration_slack = 2026-08-12 (day +1 after signup 08-11)
- Allowance set: allowance_set = 2026-08-13 (day +2)
- Admins added: admins_added = 2 (count present)
- First recognition given: first_recognition_at = 2026-08-15 14:22 (day +4)

WHAT IS NOT COMPLETE (field empty — no evidence of completion)
- HRIS integration: integration_hris is blank. 24 days elapsed at data end (08-11 to 09-04) with no connection date.
- First redemption: first_redemption_at is blank. 20 days have passed since first recognition (08-15 to 09-04) with no redemption recorded.

EARLY ENGAGEMENT SIGNALS (onboarding_usage.csv, active_givers)
- Growing, not flat: 3 on day 1 (08-11) to 15 on the last two days (09-03, 09-04). Change = 15 - 3 = +12; 12/3 = +400% (5x).
- Week 1 avg (08-11..08-17: 3,3,4,4,5,4,7): 30/7 = 4.3 givers/day. Final week avg (08-29..09-04: 11,13,11,13,13,15,15): 91/7 = 13.0. Ratio 13.0/4.29 = ~3.0x.
- Pattern: a plateau at 9 givers (08-21..08-25, except one 8) broke upward to 10-13 (08-26..09-02), then 15. Setbacks are single-day dips only (08-16, 08-18, 08-20, 08-22, 08-27, 08-31), each followed by recovery. Trajectory is positive.
- Caveats: active_givers counts recognition-giving only; it cannot show redemption behavior. No seat/employee count is provided, so % adoption (15 out of what?) cannot be computed.

THREE THINGS TO COVER ON THE CALL
1. HRIS connection. Only setup step with a blank field (integration_hris). Account otherwise moved fast — Slack day +1, allowance day +2, recognition day +4 — so this looks like an unfollowed step, not disengagement. Get it scheduled with a date.
2. Zero redemptions despite active recognition. First recognition 08-15, first_redemption_at still blank 20 days later, while givers grew to 15. With allowance_set = 2026-08-13, confirm balances actually landed in user wallets and that the reward catalog is visible — recognition with no payoff is the known stall risk.
3. Baseline the denominator and refresh the data. Pull seats/total users (not in the files) so "15 active givers" can be expressed as a %, and re-pull usage after 09-04 before the call — every conclusion above rests on a window that ends 20 days ago.
renewal-risk-conflicting-datesPASS 1.00
Plain-English question: Two systems disagree on renewal dates — which do you trust per account, and who is actually renewing inside 90 days?
Prompt — what the model saw
Produce a 90-day renewal risk brief. Two systems hold renewal dates and disagree on some accounts; decide which to trust per account and say why (multi-year contracts are known to be wrong in ChurnZero). For every renewal: company, CSM, ARR, date used, seat utilization, 3-month usage trend, risk rating with one sentence of evidence. Flag every disagreement. Close with total ARR renewing and ARR at risk.
cost $0.0601 · 247s · in 24,386 / out 19,963 tok
✓ total_renewing✓ arr_at_risk✓ disagreements_flagged✓ trust_rule

✓ Correct Answer (ground truth)

Plain-English answer: Chargebee wins per account (multi-year contracts are known wrong in ChurnZero), with each trusted date, the reason, and a 90-day-window flag.
The answer in one line: n_accounts = 20; n_disagreements = 5; snapshot_date = 2026-09-05 (+2 more below)
n_accounts20
n_disagreements5
accounts20 items — e.g. account_alias=C-0B144C78; csm=Cole Ingram; arr=30899.00; trusted_renewal_date=2026-11-02; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.4; usage_3m_ratio=1.03; risk=low; evidence=3-month usage ratio 1.03 (last3 avg 103 vs prior3 100), seat utilization 75% | account_alias=C-0B20DB64; csm=Dana Mercer; arr=21770.00; trusted_renewal_date=2026-10-07; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=56.6; usage_3m_ratio=1; risk=medium; evidence=3-month usage ratio 1.00 (last3 avg 295 vs prior3 295), seat utilization 57% | account_alias=C-0B344485; csm=Elena Sinclair; arr=64384.00; trusted_renewal_date=2026-11-16; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=78; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 241 vs prior3 231), seat utilization 78% | account_alias=C-0B369871; csm=Dana Mercer; arr=85128.00; trusted_renewal_date=2026-10-29; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.1; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 330 vs prior3 316), seat utilization 75%
arr_at_risk_90d359409.00
disagreementsaccount_alias=C-0B7D2C30; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-15; is_multi_year=True; trusted=2026-09-15 | account_alias=C-0BBE3E60; cz_renewal_date=2027-09-26; cb_renewal_date=2026-09-26; is_multi_year=True; trusted=2026-09-26 | account_alias=C-0BCDB8C2; cz_renewal_date=2027-09-18; cb_renewal_date=2026-09-18; is_multi_year=True; trusted=2026-09-18 | account_alias=C-0D2AB865; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-22; is_multi_year=True; trusted=2026-09-22 | account_alias=C-0F5D2323; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-29; is_multi_year=True; trusted=2026-09-29
snapshot_date2026-09-05
total_arr_renewing_90d1048715.00
trust_rulemulti-year contracts: Chargebee is authoritative (ChurnZero known wrong); otherwise systems agree or Chargebee wins
window2026-09-05 to 2026-12-04
graded checkstotal_renewing · arr_at_risk · disagreements_flagged · trust_rule

Model's actual answer

show the model's full answer
90-DAY RENEWAL RISK BRIEF
Basis date: 2026-09-25. 90-day window: 2026-09-25 to 2026-12-24.
Date rule: multi-year contracts are known wrong in ChurnZero, so for any CZ vs CB disagreement trust Chargebee (CB). For agreements use the agreed date.

DISAGREEMENTS FLAGGED: 5 of 20. All 5 are multi-year, all use CB date.

Seat utilization = seats_used / seats x 100.
3-month usage trend = (2026-08 active_users - 2026-05 active_users) / 2026-05 x 100.
Risk rubric applied uniformly: High = utilization <65% with trend <=-10%, OR utilization <30%; Medium = utilization 55-70% with flat trend -5% to +5%; Low = utilization >70% with stable/growing trend.

1. C-0B7D2C30 | Dana Mercer | ARR 65,901.00 | date used 2026-09-15 (CB) | DISAGREE: CZ 2026-09-10 vs CB 2026-09-15, 36mo multi-year, trust CB | seat 274/476 = 57.6% | trend (84-107)/107 = -21.5% | HIGH: steep 3-mo decline with sub-60% seats.
2. C-0BCDB8C2 | Cole Ingram | ARR 54,427.00 | date used 2026-09-18 (CB) | DISAGREE: CZ 2027-09-18 vs CB 2026-09-18, 36mo multi-year, trust CB | seat 232/424 = 54.7% | trend (110-136)/136 = -19.1% | HIGH: steep decline with sub-55% seats; CZ off by 1 year.
3. C-0D2AB865 | Elena Sinclair | ARR 38,022.00 | date used 2026-09-22 (CB) | DISAGREE: CZ 2026-09-10 vs CB 2026-09-22, 24mo multi-year, trust CB | seat 250/407 = 61.4% | trend (109-137)/137 = -20.4% | HIGH: steep decline with low-60s seats.
4. C-0BBE3E60 | Dana Mercer | ARR 30,993.00 | date used 2026-09-26 (CB) | DISAGREE: CZ 2027-09-26 vs CB 2026-09-26, 24mo multi-year, trust CB | seat 74/114 = 64.9% | trend (33-41)/41 = -19.5% | HIGH: steep decline; CZ off by 1 year.
5. C-0F5D2323 | Cole Ingram | ARR 90,647.00 | date used 2026-09-29 (CB) | DISAGREE: CZ 2026-09-10 vs CB 2026-09-29, 24mo multi-year, trust CB | seat 111/390 = 28.5% | trend (18-20)/20 = -10.0% | HIGH: lowest adoption band with double-digit decline.
6. C-0EC6999D | Elena Sinclair | ARR 79,419.00 | date used 2026-10-03 | AGREE | seat 31/112 = 27.7% | trend (15-14)/14 = +7.1% | HIGH: sub-30% seats despite flat trend.
7. C-0B20DB64 | Dana Mercer | ARR 21,770.00 | date used 2026-10-07 | AGREE | seat 214/378 = 56.6% | trend (294-296)/296 = -0.7% | MEDIUM: mid-50s seats with flat usage.
8. C-0BBC4E7A | Cole Ingram | ARR 56,374.00 | date used 2026-10-10 | AGREE | seat 228/337 = 67.7% | trend (139-142)/142 = -2.1% | MEDIUM: upper-60s seats with flat/slightly down usage.
9. C-0FD551AB | Elena Sinclair | ARR 48,815.00 | date used 2026-10-14 | AGREE | seat 210/376 = 55.9% | trend (126-125)/125 = +0.8% | MEDIUM: mid-50s seats with flat usage.
10. C-0F9F8F13 | Dana Mercer | ARR 46,230.00 | date used 2026-10-18 | AGREE | seat 199/352 = 56.5% | trend (182-182)/182 = +0.0% | MEDIUM: mid-50s seats with flat usage.
11. C-0BC34584 | Cole Ingram | ARR 16,740.00 | date used 2026-10-22 | AGREE | seat 327/494 = 66.2% | trend (106-103)/103 = +2.9% | MEDIUM: mid-60s seats with slightly up usage.
12. C-0B7A7546 | Elena Sinclair | ARR 35,062.00 | date used 2026-10-25 | AGREE | seat 182/205 = 88.8% | trend (63-61)/61 = +3.3% | LOW: near-90% seats with growing usage.
13. C-0B369871 | Dana Mercer | ARR 85,128.00 | date used 2026-10-29 | AGREE | seat 317/422 = 75.1% | trend (333-319)/319 = +4.4% | LOW: mid-70s seats with growing usage.
14. C-0B144C78 | Cole Ingram | ARR 30,899.00 | date used 2026-11-02 | AGREE | seat 169/224 = 75.4% | trend (106-99)/99 = +7.1% | LOW: mid-70s seats with growing usage.
15. C-0FC4DBB8 | Elena Sinclair | ARR 94,732.00 | date used 2026-11-05 | AGREE | seat 356/464 = 76.7% | trend (193-185)/185 = +4.3% | LOW: mid-70s seats with growing usage.
16. C-0D5BBE3A | Dana Mercer | ARR 39,740.00 | date used 2026-11-09 | AGREE | seat 85/102 = 83.3% | trend (91-87)/87 = +4.6% | LOW: mid-80s seats with growing usage.
17. C-0FB9D5AF | Cole Ingram | ARR 63,158.00 | date used 2026-11-13 | AGREE | seat 144/199 = 72.4% | trend (176-168)/168 = +4.8% | LOW: low-70s seats with growing usage.
18. C-0B344485 | Elena Sinclair | ARR 64,384.00 | date used 2026-11-16 | AGREE | seat 224/287 = 78.0% | trend (244-235)/235 = +3.8% | LOW: high-70s seats with growing usage.
19. C-0CB2C1B4 | Dana Mercer | ARR 40,628.00 | date used 2026-11-20 | AGREE | seat 386/473 = 81.6% | trend (49-50)/50 = -2.0% | LOW: low-80s seats with flat usage.
20. C-22170CA1 | Cole Ingram | ARR 45,646.00 | date used 2026-11-24 | AGREE | seat 251/294 = 85.4% | trend (146-143)/143 = +2.1% | LOW: mid-80s seats with stable usage.

Note: CB dates 2026-09-15, 2026-09-18, 2026-09-22 are past due as of 2026-09-25.

TOTALS
Total ARR renewing (all 20, all CB dates within/past-due 90-day window):
65,901 + 54,427 + 38,022 + 30,993 + 90,647 + 79,419 + 21,770 + 56,374 + 48,815 + 46,230 + 16,740 + 35,062 + 85,128 + 30,899 + 94,732 + 39,740 + 63,158 + 64,384 + 40,628 + 45,646 = 1,048,715

ARR at risk (6 HIGH): C-0B7D2C30 65,901 + C-0BCDB8C2 54,427 + C-0D2AB865 38,022 + C-0BBE3E60 30,993 + C-0F5D2323 90,647 + C-0EC6999D 79,419 = 359,409
ticket-theme-synthesis0.80
Plain-English question: What are the real themes in this quarter's support tickets, ignoring the existing tags?
Prompt — what the model saw
Synthesize the quarter's support tickets into 4 to 7 themes. Do not trust existing tags; read the text. Per theme: name, count, share, distinct accounts, ARR affected, two ticket ids, one-line recommendation. Separate broad patterns from single-account noise and rank by ARR exposure, not volume.
cost $0.0935 · 38s · in 51,112 / out 6,958 tok
✗ theme_count_range✓ counts_verifiable✓ single_account_flagged✓ arr_ranked✓ ticket_ids

✓ Correct Answer (ground truth)

Plain-English answer: 80 tickets in 4-7 themes — points-not-posting leads (20 tickets, 25%, $31,100 ARR affected) — each with counts, distinct accounts, ARR, and a recommendation.
The answer in one line: total_tickets = 80; quarter_window = 2026-06-01 through 2026-08-31
total_tickets80
noteexisting tags are noisy/unreliable; classify by text
quarter_window2026-06-01 through 2026-08-31
ranking_by_arr_exposurehris_provisioning, redemption_failures, billing_dispute, points_not_posting, slack_sync
sample_ticket_ids80 items — e.g. IC-460004, IC-460016, IC-460001, IC-460041
single_account_noise_themebilling_dispute
themes{"points_not_posting": {"count": 20, "share_pct": 25.0, "distinct_accounts": 9, "arr_affected": "31100.00", "example_ticket_ids": ["IC-460001", "IC-460002"], "single_account": false}, "slack_sync": {"
graded checkstheme_count_range · counts_verifiable · single_account_flagged · arr_ranked · ticket_ids

Model's actual answer

show the model's full answer
Total tickets: 80. Share = theme count / 80. ARR affected = sum of distinct account_alias ARR in theme. Ranked by ARR exposure.

BROAD PATTERNS:

1. HRIS provisioning / sync failure
count: 12, share: 12/80 = 15.0%, distinct accounts: 3, ARR affected: 36000 + 48000 + 30000 = 114000
accounts: C-0B2213A9, C-0DDFC9A7, C-0F6C0F34
tickets: IC-460059, IC-460060
recommendation: Fix HRIS new-hire provisioning and silent sync skips.

2. Rewards redemption / checkout failure
count: 18, share: 18/80 = 22.5%, distinct accounts: 7, ARR affected: 10300 + 10700 + 8900 + 9600 + 8700 + 9600 + 11000 = 68800
accounts: C-0B0F1BAB, C-0B827671, C-0CEF69FD, C-0D9CA315, C-0F876796, C-0FCCD2DF, C-14264ABD
tickets: IC-460025, IC-460030
recommendation: Fix checkout spins, failed gift-card delivery, and points deducted on error.

3. Points not posting / missing balances
count: 20, share: 20/80 = 25.0%, distinct accounts: 9, ARR affected: 2900 + 2700 + 4500 + 4500 + 3400 + 3500 + 4200 + 2500 + 2900 = 31100
accounts: C-0B2895EF, C-0BE96399, C-0BF20542, C-0D0B047C, C-0D284E42, C-0D3278C7, C-0D6CC8E3, C-0DD0626C, C-21FEBCBB
tickets: IC-460004, IC-460016
recommendation: Fix recognition-to-balance posting delay / delivered-but-no-points bug.

4. Slack integration / sync / slash-command failure
count: 14, share: 14/80 = 17.5%, distinct accounts: 4, ARR affected: 4400 + 3900 + 5400 + 5200 = 18900
accounts: C-0B843542, C-0BA71F12, C-10A56B0F, C-8C2E8F00
tickets: IC-460041, IC-460047
recommendation: Fix Slack channel sync, toggle reset / disconnect, and slash-command errors.

SINGLE-ACCOUNT NOISE - NOT A BROAD PATTERN:

5. Billing seat-count / renewal-tier errors - C-0E9C27D1 only
count: 16, share: 16/80 = 20.0%, distinct accounts: 1, ARR affected: 52000
accounts: C-0E9C27D1
tickets: IC-460071, IC-460069
recommendation: Resolve repeat seat-count and renewal-tier invoice errors for C-0E9C27D1 via account-level billing fix.
marketing 5 tests
case-study-matchPASS 1.00
Plain-English question: Which three existing customers are the best case-study matches for this prospect?
Prompt — what the model saw
Given the prospect, rank the three most similar existing customers with public case studies for use as social proof. Similarity on industry, size band, use case, region. Cite the fields that drove each match. Do not name any customer without a case study and never include billing data or contact names.
cost $0.0314 · 21s · in 20,018 / out 1,492 tok
✓ top3✓ decoys_excluded✓ cites_fields✓ no_billing

✓ Correct Answer (ground truth)

Plain-English answer: Three ranked matches on industry, size, use case, and region (e.g. C-11C31562), with the higher-scoring decoy explicitly excluded for a mismatched field.
The answer in one line: prospect_alias = C-82AF3719
constraintsno_billing_data, no_contact_names, only_customers_with_has_case_study_true_named
decoys_excludedcompany_alias=C-D6217CAA; similarity_score=4; reason=no_case_study | company_alias=C-C153868F; similarity_score=4; reason=no_case_study
prospect_aliasC-82AF3719
prospect_fields_usedindustry, size_band, use_case, region
top3company_alias=C-11C31562; similarity_score=3 | company_alias=C-64171065; similarity_score=3 | company_alias=C-A13C193D; similarity_score=3
graded checkstop3 · decoys_excluded · cites_fields · no_billing

Model's actual answer

show the model's full answer
Prospect C-82AF3719: industry=Technology, size_band=Mid-Market, use_case=employee_recognition, region=NA-West

Eligible pool = has_case_study=true only (8 records). Score = 1 per matching field, max 4.

1. C-64171065 - score 3/4 - 1+1+1+0=3
   match: industry=Technology, size_band=Mid-Market, use_case=employee_recognition
   mismatch: region=NA-East vs NA-West

2. C-11C31562 - score 3/4 - 0+1+1+1=3
   match: size_band=Mid-Market, use_case=employee_recognition, region=NA-West
   mismatch: industry=Manufacturing vs Technology

3. C-A13C193D - score 3/4 - 1+1+0+1=3
   match: industry=Technology, size_band=Mid-Market, region=NA-West
   mismatch: use_case=retention vs employee_recognition

Tie-break: ranked by industry+use_case matches first.
channel-attribution-mathPASS 1.00
Plain-English question: Compute cost per SQM, cost per SQO, and pipeline per dollar for every channel over six months.
Prompt — what the model saw
Compute channel performance for the trailing 6 months. Paid: spend, SQMs, SQOs, cost per SQM, cost per SQO, SQM-to-SQO rate, pipeline amount, pipeline per dollar. Organic: volume, SQO rate, pipeline. Report a channel with spend and zero SQMs as undefined, not zero. Flag rows where SQO date precedes SQM date. Finish with a reallocation recommendation and your confidence given sample sizes.
cost $0.0056 · 112s · in 10,072 / out 7,156 tok
✓ undefined_channel✓ undefined_not_zero✓ flagged_rows✓ best_channel✓ reallocation

✓ Correct Answer (ground truth)

Plain-English answer: LinkedIn Ads at $960/SQM and $3,000/SQO, organic search at $90,000 pipeline on zero spend, paid social flagged at $18,000 with zero SQMs, and a channel winner named.
The answer in one line: best_pipeline_per_dollar_channel = paid_search
best_pipeline_per_dollar_channelpaid_search
confidence_notewebinars and referral have small SQM samples (<20); treat rates as directional.
flagged_rows_sqo_before_sqmCT-000041, CT-000044
months2026-03, 2026-04, 2026-05, 2026-06, 2026-07, 2026-08
per_channel{"linkedin_ads": {"spend_usd": 24000, "sqms": 25, "sqos": 8, "pipeline_amount": 96000, "cost_per_sqm": 960.0, "cost_per_sqo": 3000.0, "sqm_to_sqo_rate": 0.32, "pipeline_per_dollar": 4.0}, "organic_sea
reallocation_recommendationfrom=paid_social; to=paid_search; rationale=paid_social has spend but zero SQMs (cost per SQM undefined, not zero); paid_search has the highest pipeline per dollar.
undefined_channelspaid_social
graded checksundefined_channel · undefined_not_zero · flagged_rows · best_channel · reallocation

Model's actual answer

show the model's full answer
CHANNEL PERFORMANCE — TRAILING 6 MONTHS (2026-03 → 2026-08)

Data checks: 118 contact rows, no duplicate contact_keys, all sqm_dates fall inside 2026-03-01 → 2026-08-28 (matches the spend window). No rows have pipeline without an SQO date.

PAID CHANNELS

| Channel | Spend | SQMs | SQOs | Cost/SQM | Cost/SQO | SQM→SQO | Pipeline | Pipe/$ |
|---|---|---|---|---|---|---|---|---|
| paid_search | $36,000 | 40 | 18 | $900.00 | $2,000.00 | 45.0% (18/40) | $720,000 | $20.00 |
| linkedin_ads | $24,000 | 24 | 8 | $1,000.00 | $3,000.00 | 33.3% (8/24) | $96,000 | $4.00 |
| paid_social | $18,000 | 0 | 0 | UNDEFINED | UNDEFINED | UNDEFINED | $0 | UNDEFINED |
| webinars | $9,000 | 11 | 4 | $818.18 | $2,250.00 | 36.4% (4/11) | $48,000 | $5.33 |
| **Paid total** | **$87,000** | **75** | **30** | | | | **$864,000** | **$9.93** |

Arithmetic:
- Spend: paid_search 6,000×6=36,000; linkedin_ads 4,000×6=24,000; paid_social 3,000×6=18,000; webinars 1,500×6=9,000. Total = 87,000.
- paid_search cost/SQM = 36,000/40 = 900.00; cost/SQO = 36,000/18 = 2,000.00; pipe/$ = 720,000/36,000 = 20.00 (18 SQOs × $40,000 each).
- linkedin_ads cost/SQM = 24,000/24 = 1,000.00; cost/SQO = 24,000/8 = 3,000.00; pipe/$ = 96,000/24,000 = 4.00 (8 × $12,000).
- webinars cost/SQM = 9,000/11 = 818.18; cost/SQO = 9,000/4 = 2,250.00; pipe/$ = 48,000/9,000 = 5.33 (4 × $12,000).
- paid_social: $18,000 spend, zero SQMs → per your rule all ratios reported UNDEFINED (division by zero), not $0.

ORGANIC / UNSPENDED CHANNELS (no spend rows exist)

| Channel | SQMs | SQOs | SQO rate | Pipeline |
|---|---|---|---|---|
| organic_search | 28 | 9 | 32.1% (9/28) | $81,000 (9 × $9,000) |
| referral | 15 | 6 | 40.0% (6/15) | $48,000 (6 × $8,000) |

MISSING DATA: organic_search and referral have no spend rows in channel_spend.csv, so cost metrics and pipeline-per-dollar cannot be computed for them. referral also has no entry in the paid set — flagged, not assumed free (no spend record exists).

DATE-INTEGRITY FLAGS (SQO date precedes SQM date) — both in linkedin_ads:
- CT-000044: SQM 2026-07-23, SQO 2026-07-18 (SQO 5 days BEFORE SQM), pipeline $12,000
- CT-000041: SQM 2026-06-14, SQO 2026-06-09 (SQO 5 days BEFORE SQM), pipeline $12,000

Sensitivity for linkedin_ads if both rows are excluded: SQOs 8→6, SQM→SQO 33.3%→6/24=25.0%, cost/SQO 24,000/6=$4,000, pipe $96,000→$72,000, pipe/$ = 72,000/24,000 = $3.00.

REALLOCATION RECOMMENDATION

1. Kill or audit paid_social immediately. $18,000 over 6 months with zero recorded SQMs — either the channel produces nothing or tracking is broken. Either way it's $18k (20.7% of paid spend) with no attributable return. First verify the tracking tag/first-touch attribution is not simply misfiring; if tracking is confirmed intact, cut it.
2. Shift that $18,000/6mo toward paid_search. It is the only clear efficiency outlier: $20 pipeline per dollar vs $4–5.33 for linkedin_ads/webinars, the best cost per SQO ($2,000) and best SQM→SQO rate (45%). Even at its current $36k, it generated $720k of the $864k paid pipeline (83.3% = 720,000/864,000).
3. Hold webinars flat, don't scale yet. $5.33 pipe/$ is respectable but only 4 SQOs / 11 SQMs — too thin to justify more budget confidently.
4. Fix linkedin_ads data before making a call on it. It's the #2 spend ($24k) with the worst pipe/$ ($4.00), and both integrity flags are its rows. On clean data it may be worse than reported ($3.00 pipe/$ under the exclusion scenario). Investigate the attribution order (likely wrong first-touch dates), then revisit.
5. Do not shift toward organic_search/referral on ROI grounds — the efficiency looks good (32–40% SQO rates, $129k pipeline combined with zero recorded spend), but with no spend data their true cost per dollar is unknowable; they're plausibly already getting effort you can't see.

Confidence: LOW-to-MODERATE.
- paid_search: MODERATE — 40 SQMs / 18 SQOs is the only channel with a sample large enough for its 45% rate to be meaningful (±~15pp at 95% confidence, roughly).
- linkedin_ads: LOW — 24 SQMs, and 2 of 8 SQOs carry date-integrity violations that swing its headline metrics by ~25–33% depending on treatment.
- webinars: LOW — 11 SQMs, 4 SQOs. A single SQO moves cost/SQO from $2,250 to $1,800–$3,000.
- paid_social: MODERATE confidence in the "zero output" fact itself, LOW confidence it reflects reality vs. a tracking failure — hence verify-before-cut.
- organic/referral: cannot be evaluated on efficiency at all (no spend data).
- The $40,000/$12,000/$9,000/$8,000 pipeline values are suspiciously uniform per channel, suggesting an assumed per-stage deal size rather than actual amounts — treat all pipe/$ figures as directional.
competitor-battlecard-updatePASS 1.00
Plain-English question: Update the Rivally battlecard with sourced pricing and where each side wins.
Prompt — what the model saw
Update the battlecard for the competitor. Sections: one-line positioning, pricing with source and date (newer source wins, note the conflict), where they win, where we win, objections and responses, recent changes, our 12-month win/loss record against them. Cite a snippet id for every factual claim. Rep opinion on a call is not a fact about the competitor. Anything from the old card you cannot re-source gets marked unverified.
cost $0.0051 · 142s · in 1,784 / out 9,639 tok
✓ competitor_named✓ pricing✓ win_loss✓ unverified_marked✓ citations

✓ Correct Answer (ground truth)

Plain-English answer: Rivally at $7 per user per month (the August source wins over the stale $5, conflict noted), a 13-7 win/loss record, and two unverified old claims flagged.
The answer in one line: competitor = Rivally
citation_ids_available25 items — e.g. S01, S02, S03, S04
competitorRivally
pricingvalue_per_user_per_month=7; source_snippet=S17; source_date=2026-08-12
rep_opinion_snippet_ids_not_factsS09, S21
unverified_old_claimsRivally lacks a Slack integration., Rivally was acquired by WorkHuman in 2025.
win_loss_12mowins=13; losses=7
graded checkscompetitor_named · pricing · win_loss · unverified_marked · citations

Model's actual answer

show the model's full answer
BATTLECARD UPDATE — Rivally (as of 2026-09-25)

POSITIONING (one line)
Points-based recognition feed (S02) that wins on engagement and EU reach but loses on analytics depth and enterprise admin tooling; sells across mid-market (S04) and enterprise (S10, S12).

PRICING
- Current list: $7/user/mo, Recognition Starter, annual billing required — pricing_page S17, 2026-08-12. NEWER SOURCE WINS.
- Conflict noted: pricing page showed $5/user/mo on 2026-01-20 (S03) and still on 2026-04-01 (S08); superseded by S17. Increase = $7 − $5 = $2 per user/mo = 2/5 = +40%.
- No pricing-page snapshot exists between 2026-04-01 (S08) and 2026-08-12 (S17), so the exact change date is between those.
- Quoted prices (call_notes deal mentions — reported quotes, not verified list):
  - $6.50/user/mo to a 500-seat prospect, annual term — S13, 2026-06-02. Doesn't match either page snapshot ($5 Apr, $7 Aug).
  - $7 list with 15% discount offered for a 3-year term — S18, 2026-08-14. Effective: 7 × 0.85 = $5.95/user/mo.
- "Rivally Pulse" engagement survey: priced as a paid add-on, not bundled — S23, 2026-09-01.

WHERE THEY WIN
- Recognition feed engagement, repeatedly praised — S02 (2025-12-15), S16 (2026-07-19).
- Fast setup; Slack integration worked out of the box — S04 (2026-02-02).
- EU/distributed teams: multi-language support praised — S12 (2026-05-21); EU data residency GA + Dublin office — S15 (2026-07-01); EU residency pitched early — S05 (2026-02-18); ex-Workday VP hired to lead EMEA — S11 (2026-05-09).
- Support response under 4 hours — S22 (2026-08-30).

WHERE WE WIN (their documented gaps; our-side superiority is only evidenced where a snippet shows it)
- Analytics: dashboards "basic compared to enterprise tools" — S07 (2026-03-22); "limited analytics" — S02. Direct win on this axis: 800-seat prospect picked Bonusly over Rivally citing analytics depth — S25 (2026-09-03, deal mention).
- Enterprise admin: no SCIM, manual user management painful — S10 (2026-04-28); no bulk recognition editing — S24 (2026-09-02); "admin tooling lags peers" — S16.
- Switching friction (use preemptively in competitive accounts): migration off Rivally hard, analytics exports CSV-only — S20 (2026-08-25).
- EMEA rewards catalog thinner than US — S14 (2026-06-14).

OBJECTIONS AND RESPONSES
1. "Rivally is cheaper." Response: current list is $7/user/mo (S17); any $5 quote is stale (S03, S08). Their real-world discount requires a 3-year term ($5.95 effective, S18). GAP: our pricing is not in the dataset — no TCO delta can be stated.
2. "Rivally has EU data residency." True as reported: GA per S15, pitched per S05. GAP: no our-side residency facts in the dataset; response needs product/legal proof before it goes on a card.
3. "Rivally sets up faster." Supported only by one mid-market review (S04). GAP: no our-side deployment data.
4. "Rivally's feed is more engaging." Supported by S02, S16. Only grounded counter: S25 — analytics depth won an 800-seat head-to-head.
5. "Rivally support is faster." Supported by S22. GAP: no our-side SLA data.

RECENT CHANGES (chronological)
- 2025-11-04: $40M Series C led by Northgate Ventures — S01.
- 2026-03-05: "Rivally Pulse" survey add-on launched — S06; 2026-09-01: exits beta, priced as add-on — S23.
- 2026-05-09: ex-Workday VP EMEA hired to lead European expansion — S11.
- 2026-07-01: Dublin office; EU data residency GA — S15.
- 2026-08-12: Starter price raised $5 → $7 (S08 vs S17; +40%).
- 2026-08-20: Microsoft Teams app v2 in public preview — S19 (preview, not GA).

12-MONTH WIN/LOSS RECORD vs RIVALLY
Scope assumption: deals file spans 2025-09 through 2026-08 = exactly 12 calendar months; all 20 rows included (month granularity only).
Wins (13): Deal-072E31, Deal-A9FD43, Deal-F65C8F, Deal-7AA785, Deal-44C524, Deal-0D0CD6, Deal-E46EAB, Deal-D5B790, Deal-1D2392, Deal-5C636E, Deal-67BE14, Deal-1B6969, Deal-F03E7B.
Losses (7): Deal-7767F5, Deal-D263E0, Deal-935746, Deal-9066A6, Deal-5645A5, Deal-72A02F, Deal-C6FFAA.
Arithmetic: 13 + 7 = 20 deals; win rate 13/20 = 65.0%.
By third: 2025-09→12 = 5W–3L (5/8 = 62.5%); 2026-01→04 = 5W–3L (5/8 = 62.5%); 2026-05→08 = 3W–1L (3/4 = 75.0%).
Caveat: the deals file has no outcome reasons. No snippet can be linked to any deal alias, so win/loss drivers per deal are MISSING DATA. S25 is the only sourced head-to-head outcome anecdote in the corpus.

OLD-CARD AUDIT
- "Points-based recognition" → re-sourced: S02.
- "for mid-market" → partially re-sourced: mid-market review S04, but enterprise reviewers S10/S12; reworded above.
- "Pricing starts at $5/user/mo (as of 2026-01)" → superseded: S03/S08 valid then, current list is $7 per S17.
- "Rivally lacks a Slack integration" → REFUTED: Slack integration worked out of the box per S04 (2026-02-02). Remove.
- "Acquired by WorkHuman in 2025" → UNVERIFIED: no snippet mentions WorkHuman or any acquisition; S01 (2025-11-04) reports an independent Series C with no acquisition reference, and S11's connection is to Workday, not WorkHuman. Remove until sourced.
- "Strong in EU enterprise with multi-language support" → re-sourced: S12, S15.

EXCLUSIONS AND RULES APPLIED
- S09 (AE opinion, "UI feels clunky") and S21 (AE opinion, "discounting aggressively") excluded — rep opinion is not a competitor fact. S18's documented 15% 3-year offer (S18) is a deal mention and is cited as a report, not verified pricing.
- All call_notes items (S05, S13, S18, S25) are cited as rep/prospect-reported events with their snippet IDs.

MISSING DATA (stated explicitly)
- Our own pricing, deployment times, SLAs, EU-residency posture, and rewards-catalog breadth — no our-side evidence exists in the corpus; three objection responses above are ungrounded until sourced.
- No 2026-09 deals in the file; no deal amounts, stages, segments, or loss reasons.
- Rivally company size/ARR/customer count: nothing beyond the S01 raise.
- Exact date of the $5→$7 change (only bracketed by S08 and S17).
nurture-sequence-diagnosis0.75
Plain-English question: What is broken in each email sequence — rates, overlaps, tracking errors?
Prompt — what the model saw
Review the sequences. Per sequence: sent, open rate, reply rate, meeting rate, weakest step. Identify tracking errors (opened above sent), audience overlap between sequences, and the failure mode for anything under 2 percent reply. One change per weak sequence and which to fix first. Under 350 words.
cost $0.0163 · 225s · in 32,204 / out 19,655 tok
✓ tracking_error✗ overlap_pair✓ weakest_step✓ fix_priority

✓ Correct Answer (ground truth)

Plain-English answer: Per-sequence open, reply, and meeting rates with the weakest step (all three weakest at step 3), an opened-above-sent tracking error, and audience overlap flagged.
The answer in one line: fix_priority_first = Cold Outbound - HR Leaders
all_pairwise_overlaps['Cold Outbound - HR Leaders', 'Cold Outbound - People Ops'], ['Expansion Nurture', 'New Logo Nurture']
audience_overlap_pairCold Outbound - HR Leaders, Cold Outbound - People Ops
failure_modeCold Outbound - HR Leaders=opens healthy (~37%) but replies near zero from step 1 -> body/CTA not landing, not a deliverability or subject-line problem
fix_priority_firstCold Outbound - HR Leaders
per_sequence{"New Logo Nurture": {"sent": 1386, "opened": 490, "replied": 90, "meetings": 27, "open_rate": 0.3535, "reply_rate": 0.0649, "meeting_rate": 0.0195, "weakest_step": 3}, "Expansion Nurture": {"sent": 8
tracking_errorsequence=Expansion Nurture; step=2; sent=300; opened=340; issue=opened exceeds sent
under_2pct_replyCold Outbound - HR Leaders
graded checkstracking_error · overlap_pair · weakest_step · fix_priority

Model's actual answer

show the model's full answer
Per-sequence metrics (totals across 3 steps; rates = events ÷ total sent):

NEW LOGO NURTURE — sent 500+458+428 = 1,386; opened 490 (35.4%), replied 90 (6.5%), meetings 27 (2.0%). Weakest step: 3 (open 120/428 = 28%, reply 18/428 = 4.2%).

EXPANSION NURTURE — sent 300+300+275 = 875; opened 565 (64.6%), replied 59 (6.7%), meetings 12 (1.4%). Weakest step: 3 (open 95/275 = 34.5%, reply 12/275 = 4.4%).

COLD OUTBOUND – HR LEADERS — sent 600+595+590 = 1,785; opened 545 (30.5%), replied 8 (0.45%), meetings 0 (0.0%). Weakest step: 3 (reply 1/590 = 0.17%, open 22%).

COLD OUTBOUND – PEOPLE OPS — sent 400+386+377 = 1,163; opened 340 (29.2%), replied 29 (2.5%), meetings 6 (0.5%). Weakest step: 3 (reply 6/377 = 1.6%, open 21%).

Tracking errors
1. Expansion Nurture step 2: opened 340 > sent 300 (113% open rate — impossible). Excluding that step, real opens are 130+95 = 225/575 = 39.1%, which is in line with peers; step 2's 25 replies/4 meetings also look inflated. Fix the counter before trusting the sequence.
2. Audience file vs sends: HR Leaders has 290 distinct contacts but 600 step-1 sends; People Ops 217 vs 400. Either sends are inflated or the audience file is incomplete — unresolvable from the data given.

Audience overlap (audiences.csv: 963 rows, 940 distinct contacts, 23 in two sequences)
- 21 contacts in both Cold Outbound – HR Leaders and Cold Outbound – People Ops (e.g., CT-000849, CT-001105, CT-001345). The two cold sequences are double-mailing the same people.
- 2 contacts in both New Logo Nurture and Expansion Nurture: CT-000301, CT-000624.

Failure mode under 2% reply — Cold Outbound – HR Leaders (0.45%)
Opens are healthy (40% at step 1, 30.5% overall) but replies collapse 0.83% → 0.17% and meetings are zero across 1,785 sends. The mail is being read and dismissed: wrong message/offer for the persona, not a deliverability problem — aggravated by 21 contacts getting two competing cold sequences.

One change per weak sequence
- HR Leaders: rewrite step 1 around one specific HR-leader pain with a concrete ask (and suppress the 21 overlaps).
- People Ops: replace step 3 (1.6% reply, 0.3% meeting) with a different channel/angle rather than a third email.
- Expansion Nurture: correct the step 2 open count.

Fix first: HR Leaders — the largest send volume (1,785) producing zero meetings is the biggest ongoing waste and is burning a shared audience.
weekly-marketing-goals-rollupPASS 1.00
Plain-English question: Are we on pace on the quarter's marketing goals?
Prompt — what the model saw
Produce the weekly marketing goals update for the active quarter: SQMs, SQOs, DS2s, closed-lost MIA rate, same-quarter close count, active pipeline coverage against target. For each: QTD actual, target, delta, pace (on/behind/ahead given days elapsed). One paragraph on what moved this week.
cost $0.0038 · 73s · in 6,632 / out 4,598 tok
✓ metrics_present✓ deltas✓ coverage_ratio✓ pace✓ paragraph

✓ Correct Answer (ground truth)

Plain-English answer: SQMs ahead of pace, SQOs on pace, DS2s behind (35 short of expected at day 66 of 92), with deltas and same-quarter closes.
The answer in one line: days_in_quarter = 92; days_elapsed = 66; coverage_ratio_active_pipeline = 0.75; elapsed_fraction = 0.7174 (+1 more below)
days_in_quarter92
coverage_ratio_active_pipeline0.75
days_elapsed66
elapsed_fraction0.7174
metrics{"SQMs": {"qtd_actual": 230, "target": 300, "delta": -70, "expected_at_day_elapsed": 215.22, "pace": "ahead"}, "SQOs": {"qtd_actual": 84, "target": 120, "delta": -36, "expected_at_day_elapsed": 86.09,
pace_ruleexpected = target * days_elapsed/days_in_quarter; within +/-5% of expected = on, above = ahead, below = behind; lower_better metrics invert.
this_week_paragraph_required1
graded checksmetrics_present · deltas · coverage_ratio · pace · paragraph

Model's actual answer

show the model's full answer
WEEKLY MARKETING GOALS UPDATE — Q3-2026 (2026-07-01 → 2026-09-30)
Days elapsed: 66 of 92 → 66 ÷ 92 = 71.7% of quarter gone. 26 days remain.
Pace method: linear expected-to-date = target × 0.717; required daily rate = remaining ÷ 26 days vs. actual daily rate = QTD ÷ 66 days.

```
metric              QTD actual   target    delta      % of tgt   linear exp.   pace
SQMs                230          300       -70        76.7%      215.2         AHEAD
SQOs                84           120       -36        70.0%      86.1          BEHIND
DS2s                40           75        -35        53.3%      53.8          BEHIND (severe)
closed_lost MIA     20.0%        ≤10.0%    +10.0pp    2.0x over  —             BEHIND
   rate (5÷25=0.20; lower_better)
same-quarter        10           20        -10        50.0%      14.3          BEHIND
closes
active_pipeline     $3,000,000   $4,000,000 -$1,000,000 75.0%   —             see note
```

Arithmetic detail:

- SQMs: expected at pace = 300 × 66/92 = 215.2; actual 230 > 215.2. Historic rate 230÷66 = 3.48/day; need 70÷26 = 2.69/day. Straight-line projection = 230÷0.717 ≈ 321 vs 300. Only metric clearing the line.
- SQOs: expected = 120 × 66/92 = 86.1; actual 84. Rate 84÷66 = 1.27/day vs needed 36÷26 = 1.38/day. Projection ≈ 117 — misses by ~3.
- DS2s: expected = 75 × 66/92 = 53.8; actual 40 — 25.7% below pace line, the biggest gap. Rate 40÷66 = 0.61/day vs needed 35÷26 = 1.35/day (2.2× current). Projection ≈ 56 vs 75.
- Closed-lost MIA rate: 5 ÷ 25 = 0.20 (20%) vs ≤0.10 target — running at 2× the allowed rate. To land the quarter at 10%, you'd need 5/(25+x) ≤ 0.10 with zero new MIA losses → x ≥ 25 additional non-MIA closed-losts. Effectively unreachable; target is missed QTD.
- Same-quarter closes: expected = 20 × 66/92 = 14.3; actual 10. Need 10 closes in 26 days (0.385/day) vs 0.152/day run-rate. Projection ≈ 14.
- Active pipeline coverage: $3.0M ÷ $4.0M = 75.0% of target. Note: this is a stock (point-in-time level), not a cumulative flow, so time-elapsed pace math doesn't strictly apply — it simply stands at 75% of target with 26 days left; the $1.0M gap must be filled by new net pipeline creation, not carry-over.

What moved this week — data gap, stated explicitly: the provided files contain only single QTD snapshots with no prior-week actuals or weekly deltas, so this week's movement cannot be measured and I won't invent it. What the QTD data does show is the structural story: volume generation is healthy but conversion decays down-funnel — SQM→SQO = 84÷230 = 36.5% and SQO→DS2 = 40÷84 = 47.6% (uncorrected for cohort maturity, so late-quarter SQMs drag the first ratio). Marketing is feeding the top (SQMs the only metric ahead of pace) but the SQO→DS2 handoff is where the quarter's deficits concentrate: closing the DS2 gap alone requires more than doubling the current DS2 entry rate, and the 20% MIA rate on 25 closed-losts says pipeline is dying unresponsive rather than losing competitively. Same-quarter closes at 50% of pace plus coverage at 75% of target compound the in-quarter revenue risk.
communication 5 tests
ceo-slack-compressionPASS 1.00
Plain-English question: Compress the forecast analysis into one Slack message the CEO can read in ten seconds.
Prompt — what the model saw
Compress the forecast analysis below into a single Slack message for the CEO. Rules: at most three sentences; under 80 words; lead with the decision or the so-what; keep every distinct number that matters and the one caveat that changes the conclusion; cut any sentence that does not carry a claim, a number, or a directive. No em dashes, no throat-clearing phrases.
cost $0.0012 · 29s · in 879 / out 1,642 tok
✓ max_3_sentences✓ under_80_words✓ weighted_present✓ caveat_present✓ no_em_dash

✓ Correct Answer (ground truth)

Plain-English answer: Three sentences, under 80 words, decision first, every distinct number kept plus the one caveat — all checked by deterministic text rules.
The answer in one line: weighted_number = 115976; must_keep = data-quality caveat — deleting the bad news fails
weighted_number115,976
limitmax 3 sentences, max 80 words, no em-dashes
must_keepdata-quality caveat — deleting the bad news fails
graded checksmax_3_sentences · under_80_words · weighted_present · caveat_present · no_em_dash

Model's actual answer

show the model's full answer
Q3 forecast is $115,977: $44,729 COMMIT (7 deals) plus 35% of $203,565 BEST_CASE ($71,248); 23 PIPELINE deals weight zero. One COMMIT deal (Deal-D348E1, $13,770) closes Oct 15 and is excluded, along with $28,240 BEST_CASE. Owner is blank on 85 of 86 deals, so nothing rolls up by rep.
stalled-deal-followup-emailPASS 1.00
Plain-English question: Write a follow-up email for this stalled deal — under 120 words, one prior touch, one ask.
Prompt — what the model saw
Write a follow-up email for this stalled deal. Under 120 words. Reference one specific prior touch from the engagement history, make one clear ask, and match the tone of the prior thread. No invented details.
cost $0.0014 · 38s · in 646 / out 2,164 tok
✓ word_cap✓ references_touch✓ single_ask

✓ Correct Answer (ground truth)

Plain-English answer: An email that references the July 28 milestone-awards demo, makes one clear ask, and invents nothing — Deal-0D2F7A, 31 days since last engagement.
The answer in one line: amount = 5100; word_cap = 120; days_since_last_engagement = 31; deal_alias = Deal-0D2F7A (+4 more below)
amount5,100
days_since_last_engagement31
word_cap120
allowed_facts14 items — e.g. Deal-0D2F7A, $5,100, DS3, Alex Franklin
allowed_newENG-00001, ENG-00002, ENG-00003, ENG-00004, ENG-00005
deal_aliasDeal-0D2F7A
must_referenceengagement_id=ENG-00002; date=2026-07-28; type=meeting; why=most substantive prior touch; concrete prospect reactions to milestone awards and points catalog
ownerAlex Franklin
single_cta1
stageDS3
tonewarm, concise, professional (matches prior thread)
graded checksword_cap · references_touch · single_ask

Model's actual answer

show the model's full answer
Subject: Following up on the July 28 People team demo

Hi [Name],

Checking in on the demo we ran with your People team on July 28. The VP People and your two HR admins had a strong reaction to the automated milestone awards and the points catalog, so I wanted to see where things landed internally.

I followed up on August 5 with a recap and the 150-seat pricing — did you get a chance to review it?

One ask: can we grab 15 minutes this week to walk through the pricing and answer any questions? Happy to send a couple of times.

Best,
Alex Franklin
funnel-digest-paragraphPASS 1.00
Plain-English question: Write the weekly GTM digest — one paragraph per team, every number traceable.
Prompt — what the model saw
Write the weekly GTM digest: one paragraph per team (Marketing, Sales, CS, PLG) in Bonusly brand voice. Every number must trace to the extracts. No customer contact names, no employee performance commentary, no compensation. Four paragraphs, under 400 words total.
cost $0.0410 · 33s · in 28,185 / out 1,690 tok
✓ four_paragraphs✓ word_cap✓ teams

✓ Correct Answer (ground truth)

Plain-English answer: Four paragraphs (Marketing, Sales, CS, PLG) using only the allowed numbers, no contact names, no performance commentary, under 400 words.
The answer in one line: word_cap = 400
word_cap400
allowed_numbers16 items — e.g. 2, 3, 9, 08
bannedcustomer contact names, employee performance commentary, compensation
required_paragraphsMarketing, Sales, CS, PLG
graded checksfour_paragraphs · word_cap · teams

Model's actual answer

show the model's full answer
Marketing kicked off the week of 2026-08-31 with 46 SQMs against a 52 target (46/52=88.5% attainment, gap 52-46=6), with the webinar channel contributing 18 of those SQMs (18/46=39.1% share) while paid search cost per SQM held at 150 dollars.

Sales kept momentum building with 14 SQOs converted and 9 DS2 meetings set (total 14+9=23 conversion motions), creating 310000 dollars in new pipeline alongside a same-quarter close count of 3 for the week.

CS protected the base we have built together, with 2 renewals saved this week and team NPS moving to 61, while carrying 3 open red-flag accounts heading into next week for focused follow-through.

PLG welcomed 412 new signups with activation at 31 percent (412x0.31=127.7 activated), and 38 companies reached the aha moment of 10 recognition gives (38/412=9.2% of signups), showing more teams finding belonging through recognition.
partner-digest-sectionPASS 1.00
Plain-English question: Write the partner digest so every partner in the roster appears exactly once.
Prompt — what the model saw
Write the partner digest for the period. Every partner in the roster appears once: active plays get status with numbers, quiet partners get one line. Partner-sourced pipeline count and amount must match the deal data. Nothing omitted, nothing added.
cost $0.0010 · 20s · in 691 / out 1,167 tok
✓ all_partners✓ pipeline_amount✓ quiet_one_line

✓ Correct Answer (ground truth)

Plain-English answer: Four active partners with status and numbers, four quiet partners with one line each, and partner-sourced pipeline that reconciles to the deal data.
The answer in one line: rule = every roster partner named exactly once; quiet partners get one line; no partners outside roster.
active_no_dealsWorkWell Group
partner_sourced_pipelinecount=5; amount=550000
per_partner_pipeline{"Apex Rewards Co": {"count": 2, "amount": 275000}, "HRCloud Partners": {"count": 1, "amount": 140000}, "CultureBridge": {"count": 2, "amount": 135000}}
quiet_partnersRecogniq, KudosWave, PeopleFirst Advisors, TotalPerk
rosterpartner_name=Apex Rewards Co; status=active | partner_name=HRCloud Partners; status=active | partner_name=CultureBridge; status=active | partner_name=WorkWell Group; status=active | partner_name=Recogniq; status=quiet | partner_name=KudosWave; status=quiet … (+2 more)
ruleevery roster partner named exactly once; quiet partners get one line; no partners outside roster.
graded checksall_partners · pipeline_amount · quiet_one_line

Model's actual answer

show the model's full answer
PARTNER DIGEST — 2026-08-24 to 2026-09-04

Active partners (4)

Apex Rewards Co — Active. 2 partner-sourced deals: Deal-DDAAF2 ($180,000) + Deal-2CF33E ($95,000) = $275,000. Both logged with UTM Source = Partner; both reported in DS1 per Cole Ingram. Co-webinar locked for 09-15 (Dana Mercer). Deal count and amounts match the deal data (2 of 2).

HRCloud Partners — Active. 1 partner-sourced deal: Deal-F1CDA5 = $140,000. Integration referral closed security review; opp moved to DS2 per Alex Franklin. Matches deal data (1 of 1).

CultureBridge — Active. 2 partner-sourced deals: Deal-096E1D ($60,000) + Deal-067213 ($75,000) = $135,000. Two sourced opps from a lunch-and-learn, both early stage per Farid Osman. Matches deal data (2 of 2).

WorkWell Group — Active. 0 partner-sourced deals, $0 (no WorkWell rows in the deal data). Wants to restart the joint playbook in Q4; planning call booked for 09-09 per Dana Mercer.

Quiet partners (4)

Recogniq — No activity in Slack export or deal data this period.
KudosWave — No activity in Slack export or deal data this period.
PeopleFirst Advisors — No activity in Slack export or deal data this period.
TotalPerk — No activity in Slack export or deal data this period.

Partner-sourced pipeline totals

Deals: 2 + 1 + 2 + 0 = 5 deals, all with utm_source = Partner — matches the 5 rows in the deal data exactly.
Amount: $275,000 + $140,000 + $135,000 + $0 = $550,000
Check: 180,000 + 95,000 = 275,000; 275,000 + 140,000 = 415,000; 415,000 + 60,000 = 475,000; 475,000 + 75,000 = 550,000.

Notes on data limits: Stage labels (DS1, DS2, early stage) come only from the Slack export — the deal file carries no stage column. All 8 roster partners appear once; nothing in the digest lacks a source row.
rfp-answer-from-policyPASS 1.00
Plain-English question: Answer the 10 RFP items from the policy excerpts only — no guessing.
Prompt — what the model saw
Answer the 10 questionnaire items from the policy excerpts only. Cite the excerpt id for every answer. Where the excerpts do not answer the question, mark it unanswerable and say what document would; do not guess. Tone: vendor response, direct.
cost $0.0006 · 19s · in 668 / out 392 tok
✓ answered_count✓ unanswerable_flagged✓ citations✓ no_guess

✓ Correct Answer (ground truth)

Plain-English answer: Seven answered with excerpt citations (Q1-Q7), three flagged unanswerable with the document that would answer them (Q8-Q10).
The answer in one line: answered_count = 7; rule = unanswerable items must be flagged and the document that would answer them named; do not guess.
answered_count7
answer_map{"Q1": ["P1"], "Q2": ["P1"], "Q3": ["P2"], "Q4": ["P3"], "Q5": ["P4"], "Q6": ["P5"], "Q7": ["P6"]}
ruleunanswerable items must be flagged and the document that would answer them named; do not guess.
unanswerableQ8, Q9, Q10
valid_citation_idsP1, P2, P3, P4, P5, P6, P7, P8
graded checksanswered_count · unanswerable_flagged · citations · no_guess

Model's actual answer

show the model's full answer
Q1 — Encryption at rest: Yes. AES-256 for all customer data. [P1]

Q2 — Encryption in transit: Yes. TLS 1.2 or higher. [P1]

Q3 — SOC 2: Yes. SOC 2 Type II, renewed annually; available to prospects under NDA. [P2]

Q4 — Post-termination retention: Data is retained for the duration of the contract plus 90 days, then deleted from production systems. [P3]

Q5 — Subprocessors: List published at trust.bonusly.com; customers receive 30 days notice before a new subprocessor is added. [P4]

Q6 — SAML SSO: Yes. SAML 2.0 (with SCIM provisioning) available on the Enterprise plan. [P5]

Q7 — Backups: Production databases backed up daily; backups encrypted and retained 35 days. [P6]

Q8 — ISO 27001: UNANSWERABLE from the excerpts — no mention of ISO certification. The compliance attestation list or ISO certificate page (e.g., trust.bonusly.com certifications section) would answer this. [No excerpt]

Q9 — Uptime SLA: UNANSWERABLE — the excerpts contain no SLA/uptime terms. The Master Services Agreement or Service Level Agreement document would answer this. [No excerpt]

Q10 — HIPAA BAA: UNANSWERABLE — no mention of HIPAA or BAAs in the excerpts. The legal/regulatory FAQ or a signed BAA template would answer this. [No excerpt]
ops-maintenance 5 tests
skill-file-reconciliationPASS 1.00
Plain-English question: Audit a skill set against its manifest and flag every conflict with a severity and an action.
Prompt — what the model saw
Reconcile the skill set below against its manifest. Find and report (with severity CRITICAL/WARNING/INFO and an action type TRIM_DESC, UPDATE_BODY, DELETE_SKILL, MERGE, or REVIEW): (1) any two skills whose ALWAYS-trigger phrases overlap/duplicate, (2) any circular delegation chain between skills (name it), (3) any delegation target that does not exist (dangling, name it), (4) any version conflict and which skill should survive, (5) how many manifest descriptions exceed 1,024 characters, (6) any hardcoded page ids, dates, or person names in skill bodies, (7) manifest drift in BOTH directions (files with no manifest row, and manifest rows with no file). Make one proposal per finding; rewrite nothing.
cost $0.0270 · 422s · in 73,869 / out 30,261 tok
✓ duplicate_trigger_found✓ circular_chain✓ dangling_target✓ oversized_desc_count✓ hardcoded_values✓ manifest_drift✓ severity_and_action

✓ Correct Answer (ground truth)

Plain-English answer: A severity-ranked list of TRIM_DESC / UPDATE_BODY / DELETE_SKILL / MERGE / REVIEW actions covering every rule violation found.
The answer in one line: files_audited = 14; rank = CRITICAL / WARNING / INFO with one fix action per issue
files_audited14
findingsduplicate trigger, circular delegation chain, dangling file reference, oversized description, hardcoded values, manifest drift
rankCRITICAL / WARNING / INFO with one fix action per issue
graded checksduplicate_trigger_found · circular_chain · dangling_target · oversized_desc_count · hardcoded_values · manifest_drift · severity_and_action

Model's actual answer

show the model's full answer
RECONCILIATION REPORT — SignalForge skill set (14 files) vs skill_manifest.csv
All counts recomputed from the provided text; nothing taken from the manifest on trust. Scope limit: only the 14 files provided. Targets outside them can only be judged "absent from the manifest," not proven nonexistent.

===============================================================
(1) OVERLAPPING / DUPLICATED ALWAYS-TRIGGER PHRASES
===============================================================

F1 — CRITICAL — MERGE
comms-drafter ↔ email-drafter. Near-duplicate trigger sets: both claim "write me an email," "draft a follow-up," "what should I say," "bump email," "contract nudge," and both fire on pasted-message feedback/rewrite/rating. Both cover outbound, follow-ups, post-demo recaps, pricing/contract follow-up, EOQ push, renewal/expansion, QBR follow-up, onboarding, check-in; both have an identical 3-option output format ("Recommended / Softer / Firmer"), identical review protocol (rate 1–10, rewrite, alternates), and the SAME contract-follow-up benchmark text verbatim in both bodies. Both carry an identical lane marker to deal-strategy-coach. Two "ALWAYS"-style owners for one request = nondeterministic routing on every drafting turn.
Proposal: merge into comms-drafter (the superset: adds support/Intercom, partner, rewards lanes); migrate email-drafter's Gmail-signature extraction, transcript priority order, and no-markdown rule into it; then repoint the two inbound references (deal-strategy-coach manager-email flow, comms lane marker). DELETE_SKILL email-drafter only after that repointing.

F2 — CRITICAL — TRIM_DESC
pipeline-intelligence-report ↔ weekly-pipeline-report collide on the same casual phrases. PIR: ALWAYS-trigger "pipeline update", "what's the pipeline look like", "run the pipeline report". Weekly: ALWAYS-trigger "update the pipeline", "what does pipeline look like", "do the pipeline report", "generate the pipeline report". These are the same utterances; both descriptions say "ALWAYS trigger" and both say never answer pipeline questions inline.
Proposal: trim weekly-pipeline-report's phrase list to week-specific strings ("weekly pipeline report", "run the pipeline update", "mid-month pipeline check", "SQM/SQO/DS2", "bookings MTD") and qualify PIR's to scored/tiered/full-pipeline asks.

F3 — WARNING — TRIM_DESC
next-to-close ("ALWAYS trigger for … which deals are most likely to close") ↔ deal-strategy-coach ("Also trigger when a manager or VP … asks which deals are likely to close"). Duplicate near-verbatim phrase on two skills.
Proposal: trim the deal-strategy-coach phrase to the per-rep coaching context ("what should I know about [rep]'s book of business" already covers it) so deal-level "most likely to close next" routes solely to next-to-close.

F4 — INFO — REVIEW
Three skills assert unconditional ALWAYS on every task/output: model-selection ("ALWAYS run this skill at the start of every task, without exception"), analysis-validator ("Always. No exceptions." + "Never skip — even on quick check requests"), signalforge-feedback ("absolute final step … Never skip"). They are complementary lifecycle ends (start / pre-publish / post-delivery), and claim-compressor and feedback each state the ordering — but model-selection's "before any … skill invocation begins" collides with validator/feedback's own "runs after every SignalForge output" chain, and no manifest field encodes the sequence.
Proposal: review whether model-selection's blanket "every task, without exception" should be scoped to SignalForge analysis tasks; keep one authoritative chain (model-selection → analysis → validator → compressor → feedback) stated in exactly one place.

===============================================================
(2) CIRCULAR DELEGATION CHAINS
===============================================================

F5 — INFO — REVIEW
No true execution cycle is provable from the 14 files. Named mutual pairs, all bounded by lane markers (each defers only the other's lane, so they terminate):
  - deal-strategy-coach ↔ email-drafter: coach hands manager-to-prospect drafting to email-drafter ("use the `email-drafter` skill which automatically retrieves your Gmail signature"); email-drafter hands strategy back ("use deal-strategy-coach instead").
  - comms-drafter ↔ deal-strategy-coach: same reciprocal structure ("this skill drafts, that skill diagnoses").
  - next-to-close → pipeline-intelligence-report is one-way (delegate for the full scored pipeline); PIR never invokes next-to-close. pipeline-intelligence-report → closed-lost-analysis (Mode 4) is one-way; Mode 4 documents itself as "called from pipeline-intelligence-report" and does not call back.
Proposal: review only — after the F1 merge, re-check the coach↔drafter pair so the cycle cannot form mid-run (coach → comms-drafter → "for deep strategy use deal-strategy-coach" → coach).

===============================================================
(3) DANGLING DELEGATION TARGETS
===============================================================

F6 — CRITICAL — REVIEW
Delegation targets named in bodies/descriptions with no manifest row and no provided file:
  - analysis-validator §12.4 (mandatory delegation — "always delegate to specialist skill"): bonusly-data-questions, bonusly-product-questions, bonusly-business-reporting-questions, bonusly-rewards-questions, bonusly-ppp-questions, bonusly-feature-flag-questions, bonusly-deal-desk-questions, bonusly-datadog-questions. Also §11: CUSTOMER_DATA_REFERENCE, HUBSPOT_CONNECTOR_REFERENCE, SIGNALFORGE_PRODUCT_INSIGHT_SKILL, skill-orchestrator.
  - comms-drafter and email-drafter: prospect-research-multithreading, bonusly-brand. deal-strategy-coach: prospect-research-multithreading (also a hard handoff trigger list).
  - sales-forecast and weekly-pipeline-report: bonusly-brand.
  - signalforge-claim-compressor: bonusly-brand; "caveman skill" (compared against, referenced as usable together).
  - signalforge-feedback: skill-orchestrator ("Skill registered in skill-orchestrator as a terminal step").
  - pipeline-intelligence-report: closed-lost-analysis resolves (file present); signalforge-reports targets are paths, not manifest names (see F7).
Severity rationale: analysis-validator's Hold/escalation logic explicitly requires delegating product/legal/flag verification; if those targets don't exist in this deployment, the mandatory QA gate has no fallback path defined.
Proposal: reconcile against the live skill tree — add manifest rows for any that exist outside this pack, or UPDATE_BODY the delegations to existing equivalents / remove them. (If they exist as org skills not listed here, this finding downgrades to INFO manifest-completeness.)

F7 — WARNING — REVIEW
File-path sidecar references unverifiable from the manifest (manifest tracks only SKILL.md rows):
  - pipeline-intelligence-report: /mnt/skills/organization/signalforge-reports/{SKILL.md, DESIGN-SYSTEM.md, signalforge.css, reports.html, brand-lockup.html} (MANDATORY PRE-BUILD reads), /mnt/skills/user/closed-lost-analysis/SKILL.md, deprecated references/html-spec.md.
  - sales-forecast: references/{data-sources.md, cadence.md, report-structure.md, report-template.html, TEMPLATE_README.md} — the scoring formula ("Full specs in references/data-sources.md") lives entirely in these sidecars, so the manifest row cannot confirm the skill is self-sufficient.
  - weekly-pipeline-report: references/{report-spec.md, queries.md}.
Proposal: extend the manifest to sidecar files (file, chars, lines per sidecar) so mandatory reads are auditable; confirm the two path roots (/mnt/skills/organization/… vs the local skill dir) resolve in the runtime.

===============================================================
(4) VERSION CONFLICTS — AND WHICH SURVIVES
===============================================================

F8 — WARNING — UPDATE_BODY (survivor: analysis-validator at v3.6)
analysis-validator carries three internal version identities: frontmatter-adjacent header "Version: 3.6" and footer "v3.6", but the Section 7 trail template prints "Validator: analysis-validator v3.2" — every published report would stamp a stale version. Related: the changelog lists 3.0→3.6 all dated May 9, 2026 with 3.6 ordered above 3.5/3.4/3.3/3.2/3.1 (out of sequence), and §269 refers to the stamp while §13's gate list says "G2-A through G2-E" in the Full Mode list vs "G2-A through G2-F" implied by the tree.
Proposal: v3.6 survives (header, footer, and pipeline-intelligence-report's "Analysis Validator v3.6" footer template agree 3-to-1 against the v3.2 trail line); fix the trail template to render the running version, and reorder/merge same-day changelog rows.

F9 — WARNING — UPDATE_BODY (survivor: pipeline-intelligence-report at v6)
pipeline-intelligence-report labels itself "version: v6 · May 2026" in frontmatter and body title, yet contains a "v4 Component Vocabulary" section and a v4 design system as the current component contract, plus "references/html-spec.md is deprecated as of May 2026" without stating what version superseded it. Two live version identities in one skill.
Proposal: v6 survives as the skill version; disambiguate "v4" explicitly as signalforge-reports design-system component-vocabulary version, not a skill version, or retitle the section.

===============================================================
(5) MANIFEST DESCRIPTIONS > 1,024 CHARACTERS
===============================================================

F10 — INFO — TRIM_DESC
Count of manifest rows exceeding 1,024 chars: ZERO. Arithmetic: max declared value = 1,006 (pipeline-intelligence-report AND signalforge-claim-compressor tied) → 1,024 − 1,006 = 18 chars headroom. Distribution >950: 1006, 1006, 1004, 996, 965, 962 = 6 of 14 rows within 7.4% of the cap; next values 945, 897.
Recount check: I re-derived all 14 folded descriptions from the YAML; 13 match the declared char count exactly, email-drafter computes 966 vs declared 965 (±1 whitespace).
Proposal: pre-emptively TRIM_DESC the four near-cap rows (pipeline-intelligence-report, signalforge-claim-compressor, partner-digest 1,004, comms-drafter 996) before any trigger phrase is added, so the cap isn't silently breached on the next edit.

===============================================================
(6) HARDCODED PAGE IDS / DATES / PERSON NAMES IN BODIES
===============================================================

F11 — CRITICAL — UPDATE_BODY
Conflicting hardcoded rosters (violates a rule another skill states as law). analysis-validator §12.3 defines "Core 6 AEs — Any 'full AE team' or 'Core 6' filter must include all six IDs": Bryce Harmon 119337721, Hugo Lindqvist 77260721, Dana Mercer 83155923, Alex Franklin 84342457, Cole Ingram 83155924, Gavin Porter 1520255671. pipeline-intelligence-report Phase 1 hardcodes "AE owner IDs (verified May 2026)" as 5 IDs — missing Hugo Lindqvist (77260721). Set difference: {Hugo Lindqvist}; ID mismatches on the 5 overlapping names: 0. stale-pipeline-report Phase 2 states "Never hardcode rep names or owner IDs. The AE roster changes" and "Rep list is always dynamic — built from Phase 2 owner resolution. Never hardcode." So PIR and the validator embed the exact anti-pattern a sibling skill prohibits; PIR's By-AE tab and leave-detection would silently omit one AE's book.
Proposal: single source of truth — resolve owner IDs dynamically (get_crm_objects on OWNERS, per stale-pipeline-report's pattern) in PIR, and reduce validator §12.3 to a pointer; if a static roster must remain, reconcile to the same set and date-stamp it in one place.

F12 — WARNING — REVIEW
Hardcoded IDs/paths that only work in one environment:
  - portal/org ID 1973303 embedded in deal URLs: next-to-close, pipeline-intelligence-report ("HubSpot org ID: 1973303"), stale-pipeline-report col K.
  - Confluence: cloudId 73fe98de-a4a3-4869-9f8a-bb1eeed4cf7f (partner-digest, sales-forecast, signalforge-feedback); spaceId 1958248479/RevOps, folder 2286616609 and page URLs 2286321666, 2265382925, 2236940297, 2237825028, 2239365136, 2238283777 (partner-digest); sales-forecast spaceId 2232811524, parent 2232582148; signalforge-feedback page 2295136266, parent "About SignalForge (2234417154)", Build Log 2247295002; deal-strategy-coach playbook URL .../pages/2257879045.
  - Slack: user ID <@U03QLMBL7AR> as the search anchor in partner-digest (hardcodes one person's threads into every run); #revops-team C0561C1JCPJ in stale-pipeline-report; fixed channel names in next-to-close (#deal-desk, #enterprise-chat, #sales-team-internal, #internal-revops) and sales-forecast (#rev-leaders).
  - Sheets: weekly-pipeline-report IDs 1CLZeOsElVDF_LF0ZG_t2nfwvhnZ6bpwqM_nX3WEYzcw and 1ENuaEcCuLjdKhMvp8FK3Ys1ek5Aw9ZuOZhsHJJFoB_k.
  - Bot-exclusion filter is placeholder garbage, not a value: analysis-validator G1-J canonical CTE `u.email not like '[email]'` — the literal token "[email]" matches nothing; the canonical 5-filter CTE is silently 4 filters.
Proposal: review — move all environment IDs to a per-deployment config the bodies reference; fix the G1-J filter to a real pattern (that one is a correctness bug feeding a hard gate, and its own G2-D placeholder rule would flag it).

F13 — WARNING — UPDATE_BODY
Hardcoded dates treated as live data:
  - analysis-validator: stage IDs 150582536–1175632767 + product tiers (§12.2), "Expected ranges (as of May 2026)" used as G1-B failure bounds (3,000–3,500 / 440,000–470,000 / 850–1,100 / 150–350), anchors "~452,000 / ~110,097" (self-contradicting the skill's own line "Do not use hardcoded figures" and G2-F's "Deal stage table in §12.2"), "Updated May 4, 2026" roster, "stale as of March 28, 2023", "CALL_SPOTLIGHT_BRIEF removed May 4, 2026", "CLOSEDWON_DEALS current as of May 4, 2026", G2-F example names "Dana Mercer"/"Gavin Porter".
  - model-selection: registry `last_checked: 2026-05-19` with a 14-day staleness rule — today (Sept 25, 2026) that registry is ~129 days past the self-imposed refresh window, so its Opus-4.7/Sonnet-4.6 tier recommendations are governed by its own rule to be re-checked before use.
  - closed-lost-analysis: sample stats as evidence — "In the 30-deal AI-field sample from May 2026: 10 of 10"; Interventions table: "17% of losses had explicit hold/pause", "8% stated budget … 14%+ had budget concern"; named examples Softheon, Estee Lauder, LIFTOFF, Nestlé, Ozinga, MinIO, Aurora Innovation, GCash, Ethos Cannabis, StickerYou; "field confirmed May 2026" constants.
  - partner-digest: "canonical reference May 16, 2026 issue", example titles "Week of May 19, 2026 / June 2, 2026", changelog 2026-05-17.
Proposal: UPDATE_BODY — move dated stats/anchors into an "as-of" block the skill re-derives at runtime (stale-pipeline-report already models this with its "run at analysis time" block); the May-4-dated "structural error" claims (CALL_SPOTLIGHT_BRIEF) need a verify-at-runtime instruction, not a hard HOLD on a date.

F14 — WARNING — TRIM_DESC / UPDATE_BODY
Person names hardcoded in routing logic (breaks silently on departure/role change):
  - Descriptions: pipeline-intelligence-report "Also trigger when Alaina or any VP asks"; stale-pipeline-report "when any AE or Alaina asks"; partner-digest "Amani's threads".
  - Bodies: sales-forecast "Manager Forecast (Alaina / VP Sales view)" and search of "#rev-leaders"; weekly-pipeline-report titles the whole skill to "Ben Lavin" and addresses Ben in Step 3/Tools ("deliver the HTML file to Ben"); partner-digest "Owner: Amani Phipps", search `from:<@U03QLMBL7AR>`; analysis-validator §12.3 roster (13 named people with IDs incl. Amani Phipps 210200121) and escalation "Finance (Manish or Amani)" in G1-K; deal-strategy-coach ICP "routed to Perseus" / ".edu routed to Farid"; stale-pipeline-report exclusion "Bonusly Support 55483190"; signalforge-feedback example titles "Gavin Porter Rep Diagnostic", "Lowe's Conversation Analysis".
Proposal: parameterize person references to role+dynamic lookup (validator already ships the lookup tool it ignores: "Resolve via HubSpot:search_owners"); strip named individuals from descriptions so triggering keys off intent, not identity.

F15 — WARNING — UPDATE_BODY
sales-forecast and weekly-pipeline-report hardcode a quarter the skill claims to generalize. sales-forecast changelog v1.1 asserts "Quarter-agnostic (Q2 → current quarter throughout)" but the body still reads "1A — HubSpot: Open Q2 Deals" and Tab 6 "Q2 Narrative" (the title template even self-contradicts: "always use the active quarter, not Q2"). weekly-pipeline-report Step 0 hardcodes "Business days complete in Q2 (April 1 – June 30, 2026; total ≈ 64–65)" — a fixed-date quarter that is now in the past for any current run — and Step 2A/2B pull "current month and remaining Q2 months"; its Q1 2026 block is declared "static": 365,152 / 475,000 = 76.9% ("77%" ✓) and 2,490,532 / 3,288,000 = 75.7% ("76%" ✓ — arithmetic checks, but these are frozen actuals inside a report that says "Always read live — do not hard-code values").
Proposal: UPDATE_BODY — complete the quarter-agnostic sweep in sales-forecast (2 literals) and parameterize weekly-pipeline-report's quarter window from today's date; drop the static Q1 block or label it as illustrative sample output.

===============================================================
(7) MANIFEST DRIFT, BOTH DIRECTIONS
===============================================================

F16 — INFO — REVIEW
Row→file: all 14 manifest rows have a corresponding .SKILL.md in the provided set (files-without-row = ∅, rows-without-file = ∅; declared_name equals file name for all 14). No orphan rows, no unmanifested skill files. Residual caveats: (a) the sidecar files of F7 (references/*.md, DESIGN-SYSTEM.md, signalforge.css, reports.html, brand-lockup.html) are invisible to this manifest — drift there cannot be assessed; (b) the email-drafter ±1 char recount (966 vs 965) is a manifest-precision nit, not drift.
Proposal: review only — extend manifest coverage to sidecars (see F7); no deletions warranted by drift.

===============================================================
SUMMARY
===============================================================
CRITICAL 4  (F1 MERGE, F2 TRIM_DESC, F6 REVIEW, F11 UPDATE_BODY)
WARNING 8   (F3, F7, F8, F9, F12, F13, F14, F15)
INFO 4      (F4, F5, F10, F16)
Q5 answer: 0 manifest descriptions exceed 1,024 chars (max 1,006; headroom 18).
Circular delegation: none provable; two bounded mutual-handoff pairs named in F5.
Dangling targets: 15 named skills/references + 12 path sidecars (F6/F7) — all relative to the 14-file manifest only; verify against the live tree before deleting anything.
No files rewritten, per instruction.
runbook-from-slack-threadPASS 1.00
Plain-English question: Turn the resolved incident thread into a runbook.
Prompt — what the model saw
Turn the resolved incident thread into a runbook: ordered steps with the exact command or action taken, who ran it, how success was verified, and rollback for each step that changed state. Every step must trace to a message in the thread. Add nothing the thread does not contain; where the thread is unclear, mark the step as needing confirmation.
cost $0.0036 · 56s · in 10,089 / out 2,971 tok
✓ ordered_steps✓ rollback✓ needs_confirmation✓ sourced

✓ Correct Answer (ground truth)

Plain-English answer: Ordered steps from the thread, each with the exact command, who ran it, how success was verified, and rollback where state changed.
The answer in one line: rule = no steps beyond what the thread contains; every step traces to a message id.
ambiguous_steporder=4; msg=M05; mark=needs-confirmation
ruleno steps beyond what the thread contains; every step traces to a message id.
stepsorder=1; msg=M01; actor=Bryce Harmon; action=acknowledge PagerDuty alert, take IC; state_change=False; verified_by=M01 | order=2; msg=M02; actor=Farid Osman; action=check queue depth; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M02 (48,213 pending) | order=3; msg=M04; actor=Farid Osman; action=pause auto-enqueue via feature flag; command=bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'; state_change=True; rollback=bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'; verified_by=M07 (queue depth falling) | order=4; msg=M05; actor=Elena Sinclair; action=cleared the dead set; state_change=True; needs_confirmation=thread gives no command or method for clearing the dead set | order=5; msg=M06; actor=Bryce Harmon; action=scale reward-worker to 6 replicas; command=kubectl scale deployment/reward-worker --replicas=6; state_change=True; rollback=kubectl scale deployment/reward-worker --replicas=3; verified_by=M07/M08 | order=6; msg=M08; actor=Cole Ingram; action=verify queue drained and error rate baseline; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M08 (returns 0) … (+2 more)
graded checksordered_steps · rollback · needs_confirmation · sourced

Model's actual answer

show the model's full answer
RUNBOOK — reward-worker queue backlog (source: #incident-2026-08-29-reward-queue, resolved)
Every step cites its thread message. Items the thread does not record are marked NEEDS CONFIRMATION — do not treat them as validated procedure.

STEP 1 — Acknowledge alert and designate incident commander [M01]
- Action: PagerDuty alert fired for reward-worker queue depth > 10k; acknowledged and took IC. (No command recorded.)
- Who: Bryce Harmon
- Success verification: None recorded; no independent verification or reversal documented.
- State change: No.

STEP 2 — Measure queue depth [M02]
- Command: `bundle exec rake sidekiq:queue_depth`
- Who: Farid Osman
- Result as reported: reward queue at 48,213 pending jobs; normal is under 500.
- State change: No (read-only).

STEP 3 — Inspect dead set [M03]
- Action: Checked dead set. Result as reported: 112 jobs, all Redis::TimeoutError from around 13:58.
- Who: Farid Osman
- NEEDS CONFIRMATION: The thread does not record the command used to inspect the dead set.
- State change: No.

STEP 4 — Pause enqueue (stop the bleed) [M04]
- Command: `bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'`
- Who: Farid Osman
- Rollback (explicitly documented in the same message): `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`
- Success verification: No direct verification of the flag state is recorded in the thread. The later queue decline [M07] and recovery [M08] are not isolated proof this action caused them.
- State change: Yes.

STEP 5 — Clear the dead set [M05]
- Action as reported: "While I was in the console I cleared out the dead set."
- Who: Elena Sinclair
- NEEDS CONFIRMATION — command: the exact command used is not recorded.
- NEEDS CONFIRMATION — verification: no verification that the clear succeeded is recorded, and no independent confirmation by another participant.
- NEEDS CONFIRMATION — rollback: none documented. This is a destructive, state-changing action; the loss of the 112 dead jobs [M03] is irreversible as far as the thread shows. This step must NOT be adopted as an approved repeatable remediation until command, verification, and rollback/data-loss posture are confirmed.
- State change: Yes.

STEP 6 — Scale workers up [M06]
- Command: `kubectl scale deployment/reward-worker --replicas=6` (was 3)
- Who: Bryce Harmon
- Rollback (explicitly documented in the same message): `kubectl scale deployment/reward-worker --replicas=3`
- Success verification: No direct verification of the replica count after the scale-up is recorded. [M07] and [M08] are later queue observations, not isolated proof of this action's effect.
- State change: Yes.

STEP 7 — Interim progress check [M07]
- Result as reported: queue depth down to 9,400, falling ~1,200/min.
- Who: Farid Osman
- NEEDS CONFIRMATION: the measurement command is not stated (it may have been the [M02] rake task, but the thread does not say so).
- State change: No.

STEP 8 — Verify recovery [M08]
- Command: `bundle exec rake sidekiq:queue_depth` — returns 0.
- Who: Cole Ingram
- Additional verification as reported: error rate in Datadog back to baseline.
- State change: No (read-only verification).

STEP 9 — Re-enable enqueue [M09]
- Command: `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`
- Who: Bryce Harmon
- Success verification (recorded in the same message): 40 new jobs processed cleanly in the next 3 minutes.
- Rollback: Not documented for this step — needs confirmation. Do not assume re-disabling the flag is the approved reversal; [M04]'s rollback text applies to [M04], not to this step.
- State change: Yes.

STEP 10 — Scale workers back down / close incident [M10]
- Command: `kubectl scale deployment/reward-worker --replicas=3`
- Who: Bryce Harmon
- Success verification: "Queue stable at 0. Incident resolved."
- Rollback: Not documented for this step — needs confirmation (the thread does not authorize scaling back to 6 as a reversal).
- State change: Yes.

Notes on scope
- Sequence in the thread: pause enqueue [M04] → clear dead set [M05] → scale up [M06] → verify empty [M08] → re-enable enqueue [M09] → scale down [M10]. The runbook preserves that order.
- No arithmetic performed per source rules; all figures (48,213 / 500 / 112 / 9,400 / ~1,200/min / 6 vs 3 replicas / 40 jobs in 3 min / timestamps) are reproduced exactly as reported.
- Open gaps needing confirmation before this becomes a repeatable runbook: dead-set inspection command [M03]; dead-set clear command, verification, and rollback [M05]; interim measurement command [M07]; rollback procedures for [M09] and [M10].
datadog-incident-timelinePASS 1.00
Plain-English question: Sequence the failure from the logs — what broke first, and what cascaded after it?
Prompt — what the model saw
From the log slice, sequence the failure: first error with timestamp, the cascade in order, the service and job involved, and the Datadog query that would confirm the first error. State what the logs do not show.
cost $0.0046 · 114s · in 3,525 / out 7,212 tok
✓ first_error_ts✓ service_named✓ cascade_ordered✓ query✓ gaps_stated

✓ Correct Answer (ground truth)

Plain-English answer: A reward-service Redis timeout at 14:01:12Z cascades to sidekiq at 14:01:40 and api-gateway at 14:03:05, with the Datadog query that confirms the first error.
The answer in one line: job_class_involved = RewardGiveJob
cascade_orderservice=reward-service; first_error_at=2026-09-03T14:01:12Z | service=sidekiq; first_error_at=2026-09-03T14:01:40Z | service=api-gateway; first_error_at=2026-09-03T14:03:05Z | service=web-app; first_error_at=2026-09-03T14:03:30Z
clean_servicespostgres
confirming_queryservice:reward-service status:error "@message:Redis::TimeoutError"
first_errortimestamp=2026-09-03T14:01:12Z; service=reward-service; message=Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s
job_class_involvedRewardGiveJob
logs_do_not_showroot cause of the Redis timeout, redis-server's own logs, infrastructure metrics (CPU/memory/network), any deploy or config change preceding the incident, who applied mitigation or when (no actor fields in logs)
graded checksfirst_error_ts · service_named · cascade_ordered · query · gaps_stated

Model's actual answer

show the model's full answer
Sequenced from the two files only (incident-response/reward-worker-queue-incident-runbook conventions applied: exact numbers, no inference beyond the slice). Note: this is the 2026-09-03 slice, not the incident-2026-08-29 thread, so none of that thread's actor names or remediation steps apply here.

FIRST ERROR
2026-09-03T14:01:12Z — reward-service, ERROR:
"Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s"

CASCADE (chronological, exact timestamps)
1. 14:01:12Z — reward-service: Redis connection timeout to redis-primary:6379 after 5s (root symptom)
2. 14:01:20Z / 14:01:30Z / 14:01:40Z — reward-service ×3: "Redis::TimeoutError: retry exhausted for RewardGiveJob"
3. 14:01:40Z — sidekiq: "RewardGiveJob failed: Redis::TimeoutError; retrying in 60s" (queue worker starts failing)
4. 14:01:46Z–14:02:58Z — sidekiq_jobs.csv: 12 RewardGiveJob failure records (J-00001…J-00012); earliest job failure J-00005 at 14:01:46Z, 34s after the first log error
5. 14:02:28Z, 14:03:31Z, 14:04:22Z, 14:05:26Z, 14:06:47Z — sidekiq: repeated "RewardGiveJob failed… retrying" (5 more, 6 sidekiq ERROR lines total)
6. 14:02:30Z — sidekiq WARN: "Queue reward depth above 10,000" (backlog signal)
7. 14:02:36Z–14:05:50Z — sidekiq_jobs.csv: 4 RecognitionDigestJob failures (J-00013…J-00016) — same error, second job class affected
8. 14:03:05Z — api-gateway: first user-facing impact, "502 upstream timeout calling reward-service /gives" (1 min 53 s after root symptom)
9. 14:03:30Z — web-app: "Give form submission failed: upstream 502 from api-gateway" (failure surfaces to end users)
10. 14:03:48Z–14:06:52Z — alternating api-gateway 502s (4 more: 14:03:48, 14:04:13, 14:05:16, 14:06:52) and web-app submission failures (3 more: 14:04:45, 14:05:42, 14:06:49)
11. 14:22:10Z — reward-service INFO: "Redis connection restored; resuming job processing"
12. 14:24:45Z — sidekiq INFO: "Queue reward depth below 500" (drain confirmed)

Cascade shape: redis-primary timeout → reward-service clients → sidekiq job retries exhausted → queue depth >10,000 → api-gateway 502 → web-app form failures. Recovery order mirrors the failure order.

SERVICES AND JOBS INVOLVED
- Services: reward-service (first error), sidekiq (queue/backlog), api-gateway (502s), web-app (user impact). redis-primary:6379 is the implicated dependency but emits no logs in this slice. postgres appears only as INFO "checkpoint complete" — no evidence of involvement.
- Jobs: RewardGiveJob (12 of 16 failure records; named in the log lines) and RecognitionDigestJob (4 records, CSV only).
- Error-line arithmetic: reward-service 4 + sidekiq 6 + api-gateway 5 + web-app 4 = 19 ERROR lines, plus 1 WARN. Job failure records: 16 (12 + 4).
- Window: first error 14:01:12Z → restore 14:22:10Z = 20 min 58 s. Last visible error 14:06:52Z; queue normalized 14:24:45Z.

DATADOG QUERY TO CONFIRM THE FIRST ERROR
Narrow window 2026-09-03 13:55 – 14:25 UTC, sorted oldest-first:

  service:reward-service status:error "Redis::TimeoutError" "redis-primary:6379"

Expect exactly one match — the 14:01:12Z "Connection to Redis at redis-primary:6379 timed out after 5s" line — with the three "retry exhausted" lines appearing only after it. A broader variant (`service:reward-service OR service:sidekiq status:error "Redis::TimeoutError"`) should show the reward-service line as the earliest hit.

WHAT THE LOGS DO NOT SHOW
- Root cause of the Redis timeout itself: no redis-primary server logs (memory, CPU, failover, network) — only client-side timeouts.
- What restored the connection at 14:22:10Z: no remediation action, actor, failover, or restart is logged in the 14:10:56Z–14:20:59Z gap (only postgres INFO checkpoints).
- Behavior between 14:06:52Z and 14:22:10Z: whether errors continued but went unlogged is unknown — the slice has no error lines in that ~15-minute span.
- Queue depth actuals: only two bracket snapshots (">10,000" at 14:02:30Z, "<500" at 14:24:45Z); peak depth, final backlog, and dead-set counts are absent.
- Log/CSV discrepancy is unexplained: 6 sidekiq "RewardGiveJob failed" log lines vs 12 RewardGiveJob failure records in sidekiq_jobs.csv; RecognitionDigestJob appears only in the CSV, never in the log slice.
- No request/trace/job IDs to link the 5 "job enqueued" INFO events or individual 502s to specific failed jobs; count of user submissions lost is unknown.
- Whether J-00001…J-00016 eventually succeeded after 14:22:10Z — the slice ends at 14:24:45Z with no per-job completion records.
feature-flag-state-summaryPASS 1.00
Plain-English question: Summarize what each feature flag controls and who has it on.
Prompt — what the model saw
Summarize the state of the listed feature flags: what each controls per the code excerpt, which companies or segments have it on, targeting rules, and any flag with no code reference or no targets. Company counts must match the export. Do not describe a flag that is not in the export.
cost $0.0013 · 32s · in 685 / out 1,975 tok
✓ orphans✓ no_targets✓ counts_match✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: Recognition streaks on for 42 beta companies, redeem-flow redesign off with a 12-company target list, plus any flag with no code reference or no targets.
The answer in one line: rule = company counts must equal the export; do not describe flags absent from the export.
code_referencedrecognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, redeem_flow_redesign, analytics_dashboard_v3, ms_teams_app_v2
flags{"recognition_streaks_v2": {"state": "on", "targeting_rules": "segment:beta_companies", "company_count": 42}, "points_budget_guardrails": {"state": "on", "targeting_rules": "all_companies", "company_c
flags_with_no_targetspaused_offboard_cleanup
orphan_flags_no_code_referencelegacy_give_modal, survey_boosters_q3, paused_offboard_cleanup
rulecompany counts must equal the export; do not describe flags absent from the export.
graded checksorphans · no_targets · counts_match · no_invented

Model's actual answer

show the model's full answer
FLAG-BY-FLAG STATE (9 flags in export; counts verbatim from flags_export.csv)

1. recognition_streaks_v2 — state: on | targeting: segment:beta_companies | companies: 42
   Per code (app/models/recognition.rb): gates StreakTracker.record(give) — i.e., recognition submissions only record streak data when enabled. On for the beta_companies segment.

2. points_budget_guardrails — state: on | targeting: all_companies | companies: 220
   Per code (app/services/budget_service.rb): gates BudgetService.new(company).enforce!(giver, points) — enforcement of point budgets per giver. On for all companies (220), the broadest flag in the export.

3. slack_dm_nudges — state: on | targeting: segment:region_na | companies: 87
   Per code (app/jobs/nudge_job.rb): job returns early unless enabled, then SlackDm.send_nudge(user) — controls whether Slack DM nudges are sent. On for the region_na segment.

4. redeem_flow_redesign — state: off | targeting: targeted_list | companies: 12
   Per code (app/controllers/redeem_controller.rb): chooses RedeemV2Component vs RedeemV1Component — the redesigned redeem flow. Flag state is off yet a targeted_list of 12 companies exists; the export does not explain how an off flag with 12 targeted companies behaves. Noted as a data observation, not interpreted.

5. analytics_dashboard_v3 — state: on | targeting: segment:tier_three | companies: 65
   Per code (app/controllers/analytics_controller.rb): sets @dashboard = AnalyticsV3.new(company) — v3 analytics dashboard. On for the tier_three segment.

6. ms_teams_app_v2 — state: off | targeting: targeted_list | companies: 9
   Per code (app/services/teams_installer.rb): gates TeamsAppV2.install(company) — installation of the v2 MS Teams app. Off, with a targeted_list of 9 companies.

FLAGS WITH NO CODE REFERENCE (3)

7. legacy_give_modal — state: off | targeting: segment:legacy_plan | companies: 14
   No reference in the code excerpt. What it controls cannot be stated from the provided data.

8. survey_boosters_q3 — state: on | targeting: segment:legacy_plan | companies: 7
   No reference in the code excerpt. Also noted: it targets the legacy_plan segment — same segment as legacy_give_modal — despite the "q3 survey boosters" name. No behavior can be described from the data.

9. paused_offboard_cleanup — state: off | targeting: (none given) | companies: 0
   No reference in the code excerpt AND no targeting rules and no companies. The only flag in the export with zero targets.

ARITHMETIC

Company counts as exported: 42 + 220 + 87 + 12 + 65 + 9 + 14 + 7 + 0 = 456 total company-flag rows.
- 42 + 220 = 262; 262 + 87 = 349; 349 + 12 = 361; 361 + 65 = 426; 426 + 9 = 435; 435 + 14 = 449; 449 + 7 = 456; 456 + 0 = 456.
- This is not a unique-company count (segments can overlap); the only lower bound on the company universe given by the export is 220 (points_budget_guardrails = all_companies).
- State tally: on = 6 flags (recognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, analytics_dashboard_v3, survey_boosters_q3, plus redeem_flow... no — see below), off = 3 (redeem_flow_redesign, ms_teams_app_v2, legacy_give_modal, paused_offboard_cleanup).
  Corrected: on = 5 (recognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, analytics_dashboard_v3, survey_boosters_q3); off = 4 (redeem_flow_redesign, ms_teams_app_v2, legacy_give_modal, paused_offboard_cleanup). 5 + 4 = 9 flags.

SUMMARY

- 6 of 9 flags have code references (flags 1–6); 3 (legacy_give_modal, survey_boosters_q3, paused_offboard_cleanup) have no code reference in the excerpt.
- Only paused_offboard_cleanup has no targeting rules and no companies.
- Data gaps: the export gives no behavior for redeem_flow_redesign and ms_teams_app_v2 beyond their code gates (both off yet with targeted lists), and nothing at all for the 3 unreferenced flags beyond their names, state, and targets.
nda-triagePASS 1.00
Plain-English question: Triage each NDA green, yellow, or red — and quote the clause that decides it.
Prompt — what the model saw
Triage each NDA as GREEN (standard approval), YELLOW (counsel review), or RED (full legal review). Cite the clause that drives each classification by quoting it. Catch embedded non-solicits, non-competes, missing carve-outs, and unusual term or governing law. One line of reasoning per NDA.
cost $0.0339 · 22s · in 19,872 / out 2,131 tok
✓ green✓ yellow✓ red✓ non_solicit_caught✓ clause_quoted✓ correct_labels

✓ Correct Answer (ground truth)

Plain-English answer: NDA-1 green (standard mutual terms), NDA-2 yellow for an embedded 18-month non-solicit, with the driving clause quoted for every classification.
The answer in one line: non_solicit_in = NDA-2
checksclassification_per_nda, clause_quotes_exist_in_document, non_solicit_caught, carveouts_checked
classifications{"NDA-1": {"classification": "GREEN", "driving_reason": "standard mutual NDA: 2-year term, standard carve-outs, Delaware law", "driving_clause_quote": "Carve-outs: information that (a) is or becomes p
non_solicit_inNDA-2
graded checksgreen · yellow · red · non_solicit_caught · clause_quoted · correct_labels

Model's actual answer

show the model's full answer
NDA-1 - [PARTY A] and Bonusly: GREEN
Driver quote: "Carve-outs: information that (a) is or becomes publicly available through no breach, (b) was known prior to disclosure, (c) is received from a third party without duty of confidence, (d) is independently developed, or (e) must be disclosed by law or court order."
Reasoning: Standard mutual NDA with all 5 carve-outs, Delaware law, and arithmetic 2 years term + 3 years survival = 5 years total obligation with no non-solicit/non-compete.

NDA-2 - [PARTY B] and Bonusly: YELLOW
Driver quote: "During the term of this Agreement and for eighteen (18) months thereafter, neither party shall solicit for employment or hire any employee of the other party with whom it came into contact under this Agreement."
Reasoning: Embedded non-solicit/no-hire requires counsel review, arithmetic 3 years term + 18 months (1.5 years) tail = 4.5 years total restriction despite standard carve-outs and Delaware law.

NDA-3 - [PARTY C] and Bonusly: RED
Driver quotes: "For a period of three (3) years following the Effective Date, Recipient shall not, directly or indirectly, engage in or provide services to any business that competes with Discloser's business." ; "This Agreement shall be governed by and construed in accordance with the laws of the Republic of Ireland, and the parties submit to the exclusive jurisdiction of its courts." ; "No confidentiality carve-outs are stated in this Agreement."
Reasoning: One-way 5-year term with 3-year non-compete plus zero carve-outs plus Ireland exclusive jurisdiction requires full legal review with no arithmetic beyond stated 5-year confidentiality / 3-year non-compete.