← back to all models ← RevenueBench home
qwen

qwen/qwen3.8-max-prime

Mean 0.953 · 32/40 perfect tests · $9.09 total · modalities in: text, image, video · out: text · each card: the prompt → the correct answer (gold) → this model's actual answer

What do the modalities mean?

Modalities in means what you can send this model: text, images, files, video, or audio. Modalities out means what it can send back. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what this model can handle in real work.

data-crm 5 tests
stage-hygiene-auditPASS 1.00
Plain-English question: How much of the open DS2+ pipeline is missing a substantive why-buys, and is anything missing DS1 timestamps or out of order?
Prompt — what the model saw
Audit stage hygiene on the open pipeline below. Rules: a deal at DS2 or later must have a substantive why-buys (15+ characters, not a placeholder such as tbd, n/a, see notes). A why_buys value of the form LONG(n) means the real text was present but truncated for display; treat it as present and substantive. Every deal should have a DS1 entry timestamp. A deal whose stage entry timestamps are out of order regressed (an earlier-stage timestamp is later than a later-stage timestamp). Report: percent of DS2+ deals non-compliant on why-buys overall and by owner (use the owners table for names), count of deals with no DS1 timestamp, count that regressed, and three process fixes. State definitions before numbers.
cost $0.9142 · 491s · in 43,782 / out 47,767 tok
✓ noncomp_pct✓ no_ds1_0✓ regressed_0✓ per_owner_breakdown✓ definitions_first✓ three_fixes

✓ Correct Answer (ground truth)

Plain-English answer: 81.5% of the 124 DS2+ deals are missing a substantive why-buys — Alex Franklin 48, Cole Ingram 19, Dana Mercer 18, Bryce Harmon 15, Farid Osman 1 — and every deal has a DS1 timestamp with none regressed.
The answer in one line: n_deals = 156; ds2plus = 124; noncomp_pct = 81.5; bryce_noncomp = 15; no_ds1 = 0; regressed = 0 (+1 more below)
n_deals156
bryce_noncomp15
ds2plus124
no_ds10
noncomp_pct81.5
regressed0
per_owner_noncompAlex Franklin=48 of 61 (78.7%); Cole Ingram=19 of 22 (86.4%); Dana Mercer=18 of 22 (81.8%); Bryce Harmon=15 of 18 (83.3%); Farid Osman=1 of 1 (100%)
graded checksnoncomp_pct · no_ds1_0 · regressed_0 · per_owner_breakdown · definitions_first · three_fixes

Model's actual answer

show the model's full answer
Stage-hygiene audit — open pipeline (156 deals)

DEFINITIONS
- DS2+ deal: current stage DS2, DS3, DS4, or DS5.
- why-buys compliant: why_buys is non-empty, ≥15 characters, not a placeholder (tbd/n-a/see notes), or of form LONG(n) with n ≥ 15 (treated as present substantive text, truncated for display). All 23 non-empty values here are LONG(n) with n ≥ 272, so all count as compliant; every other DS2+ deal has a blank why_buys.
- Missing DS1: t_ds1 empty.
- Regressed: any earlier-stage timestamp (t_dsN) strictly later than a subsequent-stage timestamp (t_dsN+1) among populated values.

RESULTS

1) DS2+ why-buys non-compliance: 101 of 124 DS2+ deals = 81.5% (arithmetic: 101/124 = 0.8145).

By owner (owners.csv names; only owners with DS2+ deals shown):
- Alex Franklin: 48/61 = 78.7%
- Dana Mercer: 18/22 = 81.8%
- Cole Ingram: 19/22 = 86.4%
- Bryce Harmon: 15/18 = 83.3%
- Farid Osman: 1/1 = 100.0% (single DS2 deal, Deal-499BF6)
Elena Sinclair has 1 deal (Deal-57FF13) but it is DS1, so excluded.

2) Deals with no DS1 timestamp: 0 of 156.

3) Deals that regressed: 0 of 156.

Data note: Deal-57FF13 (Elena Sinclair) has a short CSV row — the source field is missing; this does not affect any metric above.

PROCESS FIXES
1. Gate stage advancement on why-buys: make DS2 entry impossible without a substantive why-buys (validation rule / required-field workflow). 81.5% non-compliance means the field is optional in practice; only 23 of 124 DS2+ deals ever captured it.
2. Backfill sprint with owner accountability: route the 101 non-compliant deals to owners as tasks, prioritized by amount and forecast category (COMMIT/BEST_CASE first, e.g., Farid Osman's sole DS2 deal Deal-499BF6 is already COMMIT-tagged with no why-buys).
3. Weekly hygiene dashboard + regression monitoring: automated check on blank why-buys at DS2+, missing t_ds1, and out-of-order stage timestamps. Timestamps are clean today (0 missing, 0 regressed) — a recurring check keeps it that way and catches stage-skipping (deals entering DS3/DS4 with same-day DS2→DS3 timestamps) before it corrupts velocity analytics.
owner-resolutionPASS 1.00
Plain-English question: Every open deal has a blank owner name — resolve each one to a real person using the owners table.
Prompt — what the model saw
The deal extract below has an owner name field that is blank on nearly every row and an owner id field that is populated. Using the owners table, resolve each open deal to an owner name. Report: number of deals resolved, deals whose owner id has no match in the owners table (list them), the owner ids that map to archived/deactivated owners (list their names), and the total pipeline amount per resolved owner. Do not guess a name for any unmatched id.
cost $0.4518 · 197s · in 40,809 / out 21,650 tok
✓ resolved_count✓ unmatched_none✓ deactivated_flagged✓ jake_amount✓ justin_amount✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: All 156 deals resolve to six named owners; jake_amount maps to Bryce Harmon ($1,054,144) and justin_amount to Alex Franklin ($624,310).
The answer in one line: jake_amount = 1,054,144; justin_amount = 624,310; resolved = 156
jake_amount1,054,144
justin_amount624,310
resolved156
deactivated
per_ownerBryce Harmon=1,054,144; Alex Franklin=624,310; Dana Mercer=341,195; Cole Ingram=288,161.43; Farid Osman=4,134; Elena Sinclair=2,100
graded checksresolved_count · unmatched_none · deactivated_flagged · jake_amount · justin_amount · no_invented

Model's actual answer

show the model's full answer
Owner resolution complete. All 156 open-deal rows joined against owners.csv by owner_id.

1) DEALS RESOLVED
156 of 156 rows resolved to an owner name (every owner_id in the deals file matched exactly one row in owners.csv).

2) UNMATCHED OWNER IDS
None. No deal references an owner_id missing from the owners table. No names were guessed.

3) ARCHIVED/DEACTIVATED OWNERS
The owners table contains two archived owners (archived=true):
- 1520255671 — Gavin Porter
- 77260721 — Hugo Lindqvist
However, neither archived owner_id appears on any open deal, so zero open deals map to an archived owner.

4) TOTAL PIPELINE AMOUNT PER RESOLVED OWNER
(sum of the amount column across that owner's deals; component additions shown per owner below)

Owner (owner_id) | deals | total amount
Bryce Harmon (119337721) | 35 | 1,054,144.00
Alex Franklin (84342457)  | 67 |   624,310.00
Dana Mercer (83155923)    | 24 |   341,195.00
Cole Ingram (83155924)    | 22 |   288,161.43
Farid Osman (716654662)   |  7 |     4,134.00
Elena Sinclair (701163055)|  1 |     2,100.00
GRAND TOTAL               |156 | 2,314,044.43

Arithmetic shown:

Bryce Harmon (35 deals):
24,000 + 19,656 + 13,500 + 7,000 + 2,520 + 240,000 + 99,000 + 72,000 + 70,000 + 63,600 + 45,000 + 1 + 21,000 + 23,400 + 13,680 + 5,502 + 8,160 + 1 + 11,400 + 1 + 36,000 + 31,500 + 6,000 + 10,800 + 30,275 + 17,400 + 12,600 + 18,000 + 37,440 + 18,828 + 2,880 + 36,000 + 20,880 + 10,920 + 25,200 = 1,054,144.00

Alex Franklin (67 deals):
14,850 + 13,770 + 11,200 + 9,000 + 6,360 + 5,400 + 3,240 + 2,484 + 1,920 + 1,080 + 7,200 + 19,000 + 2,880 + 1,400 + 4,800 + 1,632 + 10,000 + 9,300 + 2,700 + 2,160 + 1,800 + 3,600 + 3,840 + 15,000 + 1,968 + 4,000 + 3,600 + 4,800 + 3,120 + 2,520 + 9,000 + 2,400 + 62,000 + 5,400 + 5,100 + 16,700 + 4,400 + 1,620 + 2,600 + 7,200 + 18,000 + 17,000 + 8,316 + 8,100 + 18,000 + 12,600 + 24,000 + 15,000 + 9,000 + 7,200 + 3,780 + 16,200 + 7,200 + 4,680 + 1,800 + 18,000 + 2,730 + 2,400 + 3,060 + 18,000 + 12,000 + 1,800 + 4,400 + 31,200 + 7,200 + 1,600 + 60,000 = 624,310.00

Dana Mercer (24 deals):
11,250 + 10,500 + 9,000 + 9,000 + 5,400 + 4,800 + 4,600 + 1,920 + 15,000 + 4,200 + 18,900 + 27,000 + 43,875 + 20,000 + 60,000 + 8,100 + 16,250 + 3,150 + 5,000 + 2,100 + 23,400 + 5,400 + 7,350 + 25,000 = 341,195.00

Cole Ingram (22 deals):
2,760 + 1,330 + 7,225.40 + 3,360 + 5,616 + 2,700 + 7,781.20 + 6,947.50 + 40,000 + 12,168 + 2,480.40 + 11,116 + 11,193 + 1,875 + 31,750 + 32,175 + 4,779.88 + 3,334.80 + 58,529.25 + 4,140 + 18,000 + 18,900 = 288,161.43

Farid Osman (7 deals):
1,249 + 1 + 1 + 1 + 2,880 + 1 + 1 = 4,134.00

Elena Sinclair (1 deal):
2,100 = 2,100.00

Notes on data quality (no invention, flagged only):
- Several deals carry amount = 1 (Deal-012CB1, Deal-483B2D, Deal-3795AD, Deal-03BA69, Deal-117863, Deal-F17780, Deal-2B39B0, Deal-8FDCD2) — likely placeholder values; included as-is per instruction to use only provided data.
- Elena Sinclair's single deal (Deal-57FF13, 2,100, close 2026-09-02) is already past its close date relative to the other dates in the file.
- Deal-3EED2C has n_contacts=2 but a blank last_contacted_field; not material to this report.
gong-hubspot-join-integrity0.75
Plain-English question: How many call records point at deal ids that do not exist in the open pipeline?
Prompt — what the model saw
Check the integrity of the call-to-deal mapping below. The gong table lists calls in the last 90 days keyed to HubSpot deal ids; deals_open lists the current open default-pipeline deals. Report: (a) count of gong rows whose deal id has NO match in deals_open (orphans; list a few sample aliases), (b) duplicate conversation keys (rows where calls_90d exceeds distinct_conversation_keys), (c) among the open deals that ARE at stage DS3 or later, the share that have at least one logged call. Give sample aliases for the orphan finding, all taken from the data. Note: gong_owner_id is a Gong user id, not a HubSpot owner id.
cost $0.5705 · 243s · in 53,065 / out 25,412 tok
✓ orphan_count✗ dup_keys_0✓ ds3_coverage✓ sample_ids_present

✓ Correct Answer (ground truth)

Plain-English answer: 30 Gong call rows are orphans keyed to dead deal ids, none are duplicates, and DS3 call coverage is 29.4%.
The answer in one line: orphans = 30; ds3_cov_pct = 29.4; dup = 0
orphans30
ds3_cov_pct29.4
dup0
orphan_aliasesDeal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42, Deal-36EA09
graded checksorphan_count · dup_keys_0 · ds3_coverage · sample_ids_present

Model's actual answer

show the model's full answer
Call-to-deal integrity check — gong_calls_by_deal_90d.csv (67 rows) vs deals_open.csv (156 rows). Join key: hs_deal_id → deal_id only (gong_owner_id ignored per instruction — it is a Gong user id).

(a) ORPHANS — gong rows with no match in deals_open: 30 of 67
    30 / 67 = 44.8% of gong rows.
    Sample aliases (all from the data):
      Deal-8FA85D (60251290957, 46 calls)
      Deal-8FC3F9 (60251649055, 24 calls)
      Deal-3B7945 (60251639682, 21 calls)
      Deal-42B265 (61227242540, 21 calls)
      Deal-9CCC42 (61430316324, 17 calls)
      Deal-3F86A0 (60419904928, 6 calls)
    Observable pattern (not an explanation — the data doesn't say why): every orphan has calls_with_ds2_date = 0 AND calls_with_sqo_date = 0, while every matched row has nonzero values in at least one of those columns. Also, 16 of the 30 orphan deal ids share the 60251... prefix. Several orphans have last_call dates after 2026-09-24 (e.g. 2026-11-27), which is outside the snapshot window implied by deals_open.

(b) DUPLICATE CONVERSATION KEYS — rows where calls_90d > distinct_conversation_keys: 0
    All 67 rows have calls_90d exactly equal to distinct_conversation_keys (checked row by row; min gap = 0). No duplication detected on this metric.

(c) DS3+ CALL COVERAGE — open deals at stage DS3, DS4, or DS5 (current stage column, not timestamps):
    DS3+ open deals:        85  (DS3 = 61, DS4 = 14, DS5 = 10)
    With ≥1 logged call:    25
    Without any call:       60
    Share = 25 / 85 = 0.294 = 29.4%

    Notable: coverage is not uniform by stage — the 25 covered deals include all high-activity matches, while 60 DS3+ deals (e.g. Deal-9AAE5F DS4, Deal-403845 DS5, Deal-2465CE DS5, Deal-FD9F4E DS5) have zero rows in the gong table. Four DS5 deals with no logged calls: Deal-403845, Deal-2465CE, Deal-FD9F4E, plus... (exactly 3 DS5 unmatched; the fourth DS5, Deal-D348E1, Deal-C26D20, Deal-547B2B, Deal-B7EBD1, Deal-A2B47C, Deal-C61CF7, Deal-584EE5 ARE covered). Correction from the run output: unmatched DS5 = Deal-403845, Deal-2465CE, Deal-FD9F4E (3 of 10 DS5 deals uncovered).

Bottom line: the mapping fails integrity on two of three checks — 44.8% orphan rate and only 29.4% call coverage of DS3+ open pipeline. Conversation-key uniqueness is clean.
snowflake-sql-trapped-catalogPASS 1.00
Plain-English question: Write one Snowflake query for first-month usage per customer using only the provided, trap-laden catalog.
Prompt — what the model saw
Using ONLY the catalog provided, write one Snowflake SQL query returning, per customer company, for its first calendar month as a customer: unique givers, recognition count, successful redemption count. Read the catalog notes carefully: some tables are stale or unpopulated and must be avoided, and one documented filter (excluding deleted givers) must NOT be applied to historical giving counts. Put a comment above each FROM or JOIN explaining the table choice. Return only fenced SQL.
cost $0.0806 · 81s · in 1,301 / out 5,469 tok
✓ avoids_stale_deals✓ avoids_unpopulated_live✓ uses_cohort_summary✓ succeeded_filter✓ no_giver_deletion_filter✓ returns_sql

✓ Correct Answer (ground truth)

Plain-English answer: A single correct SQL query that uses the right catalog tables and avoids the trap columns — graded on the exact tables and columns.
The answer in one line: use_table = PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
avoid_stalePRODUCTION.HUBSPOT.DEALS (stale)
avoid_unpopulatedV2_LIVE.OBJECTS_DEALS (unpopulated)
forbidden_filterdeleted-giver exclusion must NOT be applied
outputone fenced SELECT with per-FROM/JOIN table-choice comment
required_filtersucceeded redemptions only
use_tablePRODUCTION.PLG.COMPANY_COHORT_SUMMARY
graded checksavoids_stale_deals · avoids_unpopulated_live · uses_cohort_summary · succeeded_filter · no_giver_deletion_filter · returns_sql

Model's actual answer

show the model's full answer
```sql
-- ============================================================================
-- Goal: per customer company, for its FIRST CALENDAR MONTH as a customer:
--       unique givers, recognition count, successful redemption count.
-- Built strictly from schema_catalog.md. No facts, columns, or numbers invented.
--
-- EXPLICIT DATA GAPS (stated, not papered over):
--  1. UNIQUE GIVERS: NOT AVAILABLE. No catalog table exposes giver identity or
--     a per-company unique-giver count. M1_USERS counts users, not givers, and
--     is therefore NOT substituted. Returned as NULL.
--  2. COVERAGE: the catalog contains no recognition/giving EVENT table and no
--     deal->company association, so sales-pipeline customers cannot be measured.
--     The only per-company first-month metrics documented anywhere in the
--     catalog are the pre-aggregated M1_* columns on the PLG cohort table, so
--     results cover self-serve companies only.
--  3. M1_* window semantics (calendar month vs. first 30 days; anchored to
--     SIGNUP_DATE, ACTIVATED_DATE, or FIRST_SUB_PAYMENT_DATE) are NOT documented.
--     M1_* is used as-is because it is the only first-month measure available.
--  4. PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 is the documented
--     redemption-event source (STATE = 'succeeded' only), but the catalog lists
--     no company key and no event-date column for it, so it CANNOT be scoped per
--     company or to a first-month window without inventing schema. The
--     pre-aggregated M1_REDEMPTIONS is used instead; whether it already applies
--     STATE = 'succeeded' is not documented.
--  5. The company key column of COMPANY_COHORT_SUMMARY is not named in the
--     catalog excerpt ("one row per self-serve company"). A clearly-flagged
--     placeholder is used below rather than guessing a real column name.
--
-- TABLES DELIBERATELY AVOIDED (per catalog notes):
--  * HUBSPOT_HUB_1973303.V2_LIVE.OBJECTS_DEALS -> exists but UNPOPULATED.
--  * PRODUCTION.HUBSPOT.DEALS                  -> STALE (last sync 2023-03).
--  * PRODUCTION.HUBSPOT.GONG_HUBSPOT_MAP_FAST  -> RETIRED pattern.
--  * HS_ENGAGEMENTS_ENRICHED / GONG_* tables   -> emails, calls, meetings,
--    transcripts: engagement data, not recognition or redemption events.
--  * CHARGEBEE.SUBSCRIPTIONS                   -> billing only; no company name,
--    no giving/redemption measures.
--
-- BUSINESS RULE HONORED: the deleted-givers exclusion filter is NOT applied to
-- these historical giving counts (catalog: applying it understates history).
-- ============================================================================

SELECT
    COMPANY_ID                                   AS customer_company_id,          -- PLACEHOLDER: company key column not named in catalog excerpt; grain is one row per self-serve company
    DATE_TRUNC('month', FIRST_SUB_PAYMENT_DATE)  AS first_customer_month,         -- first calendar month as a (paying) customer; SIGNUP_DATE/ACTIVATED_DATE are pre-customer milestones
    CAST(NULL AS INTEGER)                        AS unique_givers,                -- NOT AVAILABLE in catalog (gap #1); intentionally not proxied by M1_USERS
    M1_GIVES                                     AS recognition_count,            -- documented month-1 giving metric; deleted-giver filter deliberately NOT applied (business rule)
    M1_REDEMPTIONS                               AS successful_redemption_count   -- documented month-1 redemption metric (see gap #4 re: STATE='succeeded')
-- FROM: PRODUCTION.PLG.COMPANY_COHORT_SUMMARY is the ONLY catalog table holding
-- per-company first-month give/redemption metrics plus a customer start date
-- (FIRST_SUB_PAYMENT_DATE). Stale/empty/retired alternatives are listed above.
FROM PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
WHERE FIRST_SUB_PAYMENT_DATE IS NOT NULL   -- restrict to companies that became paying customers (documented column)
;
```
crm-completeness-enrichmentPASS 1.00
Plain-English question: Audit the CRM for missing and invalid fields, and propose the fixes.
Prompt — what the model saw
Audit the CRM extract for completeness and propose fixes. Deals need owner, stage, amount, close date, why-buys; companies need industry, employee count, HQ country; contacts need email, title, persona. Report completeness percent per field. Identify duplicate company clusters (name variants, shared domain) with a survivor per cluster. List invalid emails and domain mismatches. Fill missing company fields only where the enrichment export has a matching row; where CRM and enrichment disagree list both and recommend a source. Never invent a value. End with the 10 fixes with the most pipeline amount at stake.
cost $0.1922 · 102s · in 12,170 / out 9,229 tok
✓ dup_clusters✓ invalid_emails✓ survivor_language✓ completeness_pct✓ no_invent

✓ Correct Answer (ground truth)

Plain-English answer: 9 companies missing employee count, 6 missing HQ country, 4 invalid emails, 1 domain mismatch, 2 duplicate clusters, and 25 enrichment matches.
The answer in one line: n_contacts = 52; n_companies = 34; enrichment_matches = 25; invalid_email_count = 4; n_dup_clusters = 2; domain_mismatch_count = 1 (+4 more below)
n_contacts52
domain_mismatch_count1
enrichment_matches25
invalid_email_count4
n_companies34
n_dup_clusters2
company_missingindustry=0; employee_count=9; hq_country=6
domain_mismatchesCT-0011
duplicate_clusters['C-0A092931', 'C-0A092932'], ['C-0A092933', 'C-0A092934']
invalid_emailsCT-0010, CT-0080, CT-0081, CT-0192
graded checksdup_clusters · invalid_emails · survivor_language · completeness_pct · no_invent

Model's actual answer

show the model's full answer
CRM COMPLETENESS AUDIT
=======================

DATA GAP (stated up front): No deals extract was provided. Deal fields (owner, stage, amount, close date, why-buys) cannot be audited, and no pipeline dollar amounts exist in the data. I will not invent them. The "top 10 fixes" section is therefore ranked by records/data at stake, not dollars.

1. COMPLETENESS PER FIELD
-------------------------
Companies (34 rows):
  industry        34/34  = 100.0%
  employee_count  25/34  =  73.5%   missing: C-EC3025, C-96039F, C-44EA29, C-D04904, C-B23205, C-60C75F, C-7BBDFA, C-50D386, C-93C8BF
  hq_country      28/34  =  82.4%   missing: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB

Contacts (52 rows):
  email (valid format)  48/52 = 92.3%
  title                 39/52 = 75.0%
  persona               37/52 = 71.2%

Deals: NOT PROVIDED — 0 rows to audit.

2. DUPLICATE COMPANY CLUSTERS
-----------------------------
Cluster A — shared domain acme-corp.com:
  C-0A092931 (Technology, 500, US)
  C-0A092932 (tech, 510, USA)
  Survivor: C-0A092931 — industry and country match the CRM's dominant canonical forms ("Technology", "US"). Conflict on merge: employee_count 500 vs 510; no enrichment row exists for acme-corp.com to arbitrate, so recommend manual verification before merge.

Cluster B — shared domain globex.io:
  C-0A092933 (SaaS, 200, US)
  C-0A092934 (Technology, 200, US)
  Survivor: C-0A092934 — employee_count and hq_country identical; "Technology" matches the canonical industry used across the file ("SaaS" appears nowhere else). No enrichment row for globex.io to confirm.

No other shared domains or name variants found (all remaining 30 aliases have unique domains).

3. INVALID EMAILS AND DOMAIN MISMATCHES
---------------------------------------
Invalid (malformed, no domain after @):
  CT-0010  user0@   (company C-66D1FC)
  CT-0080  user0@   (company C-92D97D)
  CT-0081  user1@   (company C-92D97D)
  CT-0192  user2@   (company C-425E2A)

Domain mismatch (valid format, wrong domain vs company):
  CT-0011  user1@other-domain.com  on company domain 66d1fc.com (C-66D1FC)

Note: CT-0010/CT-0011/CT-0012 are all "VP People / champion" at C-66D1FC — likely the same person entered three ways; only CT-0012 has a clean email. Candidate contact-level dedupe after email repair.

4. ENRICHMENT FILLS (zoominfo_enrichment.csv matched by domain)
---------------------------------------------------------------
Fills applied only where CRM blank AND ZI row has a value (8 fills, all employee_count):
  C-EC3025  employee_count  <- 400
  C-96039F  employee_count  <- 400
  C-44EA29  employee_count  <- 400
  C-D04904  employee_count  <- 400
  C-B23205  employee_count  <- 400
  C-60C75F  employee_count  <- 400
  C-7BBDFA  employee_count  <- 400
  C-50D386  employee_count  <- 400
Post-fill employee_count completeness: 33/34 = 97.1%.

Unfillable (no value invented):
  hq_country — ZI row exists but blank: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5
  hq_country — no ZI row: C-EE9FFB (ee9ffb.com)
  employee_count — no ZI row: C-93C8BF (93c8bf.com)
  No ZI row at all (9 companies): C-BA969B, C-332637, C-93C8BF, C-EE9FFB, C-C9BB20, C-0A092931, C-0A092932, C-0A092933, C-0A092934

5. CRM vs ENRICHMENT DISAGREEMENTS (both values present)
---------------------------------------------------------
Employee_count: ZERO substantive conflicts (all overlapping rows agree numerically).

The 20 remaining diffs are format variants, not true conflicts:
  hq_country (9 rows): CRM 'US'/'USA' vs ZI 'United States' — C-66D1FC, C-950043, C-EC3025, C-96039F, C-77A95A, C-B23205, C-E51FB7, C-D0662E, C-425E2A, C-2D7423
  industry (11 rows): CRM 'tech'/'Tech '/'Technology' vs ZI 'Computer Software' — C-66D1FC, C-EC3025, C-44EA29, C-92D97D, C-D04904, C-77A95A, C-AA8DDA, C-B25F40, C-60C75F, C-425E2A
Recommendation: normalize to one canonical vocabulary rather than pick a source — 'United States' for country (ZI form, ISO-consistent) and map ZI 'Computer Software' + CRM 'tech'/'Tech ' → 'Technology' (the CRM's most-used canonical label, 10 rows). Also normalize 'health care' → 'Healthcare' (C-7BBDFA, C-50D386 — CRM and ZI agree on the non-canonical spelling). Where a true conflict later appears, prefer ZI for firmographics (employee_count, hq_country) and CRM for industry segmentation, since ZI is the systematic enrichment source.

6. TOP 10 FIXES (ranked by records/data at stake — pipeline dollars UNRANKABLE, no deals file provided)
------------------------------------------------------------------------------------------------------
 1. Merge acme-corp.com cluster (C-0A092931 survivor, C-0A092932 folded in); verify 500 vs 510 headcount manually first — 2 company records, split-history risk.
 2. Merge globex.io cluster (C-0A092934 survivor) — 2 company records.
 3. Repair 4 malformed emails (CT-0010, CT-0080, CT-0081, CT-0192) — blocks all outreach/sequencing on 4 contacts incl. 2 economic-buyer-adjacent accounts (C-92D97D, C-425E2A).
 4. Resolve CT-0011 domain mismatch (user1@other-domain.com vs 66d1fc.com) and dedupe the triple "VP People/champion" entries at C-66D1FC.
 5. Apply 8 employee_count fills from ZI (=400 each) — lifts field to 97.1%; these 8 accounts include segment/ICP-filtering candidates.
 6. Backfill 13 missing contact titles — persona/targeting impossible without them (CT-0000, CT-0022, CT-0072, CT-0080, CT-0081, CT-0092, CT-0120, CT-0121, CT-0122, CT-0132, CT-0141, CT-0162, CT-0170). C-AA8DDA has all 3 contacts title-less.
 7. Backfill 15 missing personas — including economic-buyer coverage gaps (CT-0060, CT-0082 etc.); 6 accounts have zero identified buyer persona.
 8. Normalize hq_country vocabulary (US/USA/United States → United States) — 10 conflicting rows + 6 blanks; fixes reporting segmentation.
 9. Normalize industry vocabulary (tech/Tech /Technology/Computer Software → Technology; health care → Healthcare; decide SaaS mapping) — 13 rows; fixes industry-based pipeline cuts.
10. Source enrichment for the 9 companies with no ZI row (C-BA969B, C-332637, C-93C8BF, C-EE9FFB, C-C9BB20 + the 4 dup-cluster rows pre-merge) and the 5 ZI-blank hq_country rows (C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5) — 7 fields remain unfillable from the data provided.

Reminder: to produce the requested ranking by pipeline amount at stake, and to audit deal owner/stage/amount/close date/why-buys completeness, provide the deals extract. Nothing above was ranked by dollars because no dollar values exist in the inputs.
deal-intelligence 5 tests
closed-lost-classificationPASS 1.00
Plain-English question: Why did we lose the last six months of deals — one primary category per deal, and whose side is the reason on?
Prompt — what the model saw
Classify every deal lost in the last 6 months below into one primary category: pricing, competitor, no decision, timing, product gap, champion left, or other. Also assign a side (Bonusly, buyer, unknown). Then summarize: category counts, the side split, the number of deals where the structured closed_lost_tag clearly disagrees with the free-text reason, and the two patterns most worth acting on. Base everything only on the tag and free-text provided.
cost $0.2312 · 154s · in 15,628 / out 12,176 tok
✓ total_90✓ competitor_top✓ category_counts_present✓ side_split✓ two_patterns

✓ Correct Answer (ground truth)

Plain-English answer: Every closed-lost deal gets exactly one category (pricing, competitor, no decision, timing, product gap, champion left, other) plus a side, per the computed classification.
The answer in one line: total_lost_6mo = 90; top_reason = competitor (26)
total_lost_6mo90
categoriespricing, competitor, no decision, timing, product gap, champion left, other
patternstwo most actionable, tag-vs-text disagreements counted
sideBonusly / buyer / unknown per deal
top_reasoncompetitor (26)
graded checkstotal_90 · competitor_top · category_counts_present · side_split · two_patterns

Model's actual answer

show the model's full answer
Scope note: classification below uses ONLY closed_lost_tag + closed_lost_reason_text from the file provided. The ai_closed_lost_reason field and conversation evidence (per the closed-lost-analysis standard) are NOT in this file, so confidence is bounded by rep-entered text. 90 deals, all classified. "MIA/unresponsive" treated as an outcome → mapped to no decision with unknown side unless text gives a cause.

DEAL-LEVEL CLASSIFICATION (alias · category · side)

```
Deal-DB0AAC  timing        buyer     Deal-1E7DA9  competitor    buyer
Deal-F7F635  competitor    buyer     Deal-2BBA21  no decision   unknown
Deal-AC944F  no decision   unknown   Deal-286F9C  competitor    buyer
Deal-214060  no decision   unknown   Deal-7FBAC6  no decision   buyer
Deal-91A056  timing        buyer     Deal-369281  competitor    buyer (Paylocity/HRIS-native)
Deal-29326C  timing        buyer     Deal-386F6E  no decision   unknown
Deal-5DB9B0  other (spam/not-ICP) Bonusly  Deal-9FCD0D competitor buyer (Canadian co)
Deal-831B7B  timing        buyer     Deal-55867E  no decision   buyer
Deal-F97C37  competitor    buyer     Deal-DAFB82  pricing       buyer
Deal-13E9CF  no decision   buyer (deprioritized; "not a budget issue")
Deal-39E25C  timing        buyer     Deal-2FEDDB  no decision   buyer
Deal-7ED004  pricing       buyer     Deal-64B19A  competitor    buyer (Motivosity)
Deal-21B045  no decision   unknown   Deal-3F86A0  no decision   unknown
Deal-B3ABED  timing        buyer     Deal-096750  no decision   unknown
Deal-422BA6  competitor    buyer (ADP TotalSource PEO partner)
Deal-ED9AE7  other (timing+budget+authority mix) buyer
Deal-988493  no decision   unknown   Deal-F325A5  champion left buyer (layoffs + leadership change)
Deal-381C8C  no decision   unknown   Deal-ABD14C  no decision   buyer
Deal-F308CA  no decision   unknown   Deal-79E61A  no decision   unknown
Deal-F1E8A6  no decision   unknown   Deal-8A119B  pricing       buyer
Deal-B6AC09  timing        buyer     Deal-AE7C4E  no decision   unknown
Deal-70F704  no decision   unknown (scope: anniversary awards only, then MIA)
Deal-E6E80A  timing        buyer     Deal-DAB4F1  no decision   unknown
Deal-B038F0  timing        buyer     Deal-B4B50F  no decision   unknown
Deal-4664E1  no decision   unknown   Deal-981AD4  product gap   Bonusly (UI fit, not UK-focused)
Deal-175756  timing        buyer     Deal-DC77FE  competitor    buyer (customization: points-as-dollars)
Deal-E74A73  no decision   buyer (manual test first)
Deal-DDAB52  competitor    buyer (Rippl)   Deal-5885B9 no decision unknown
Deal-ACE061  competitor    buyer (HeyTaco, rep-inferred)
Deal-BB78F3  timing        buyer     Deal-F325A5 listed above
Deal-D48E0B  no decision   unknown   Deal-A2C349  competitor    buyer (Awardco incumbent)
Deal-15DA99  timing        buyer     Deal-9F176A  timing        buyer
Deal-F4AF5D  timing        buyer     Deal-7B2236  pricing       buyer (budget + "simpler and cheaper")
Deal-79B7A1  timing        buyer     Deal-AFA56C  no decision   unknown
Deal-583ADB  no decision   unknown   Deal-C7156E  competitor    buyer
Deal-8E27DA  other (swag-only scope, no R&R) buyer
Deal-2D2F8D  competitor    buyer     Deal-C33D91  pricing       buyer (budget cuts)
Deal-E0441F  other (stale handoff from departed rep) Bonusly
Deal-7CB44D  no decision   unknown   Deal-9048EB  product gap   Bonusly (bad fit + multiple feature gaps)
Deal-0F96AA  competitor    buyer (cut before RFP finalists)
Deal-1BCA50  competitor    buyer (stakeholder committed to other vendor; text says budget-primary)
Deal-7CC678  competitor    unknown ("Nothing specific provided" — tag only)
Deal-FAC17C  no decision   buyer (contract out 2 months, Exec IT Director approval never came)
Deal-242273  competitor    buyer (points-currency digitization differentiator)
Deal-50E5D8  no decision   buyer (leadership pause)
Deal-5E64CE  competitor    buyer (Nectar contract locked to Oct 2027, exit fee high)
Deal-8A0992  competitor    buyer (Canadian provider)
Deal-D0C698  competitor    buyer (Kudos, past user)
Deal-69CF3D  timing        buyer (On Hold)
Deal-ECBF89  timing        buyer (On Hold)
Deal-3618CC  product gap   Bonusly (wanted Surveys)
Deal-EECC02  competitor    buyer
Deal-5AD03E  product gap   Bonusly ("more defined budget access")
Deal-D1A623  timing        buyer
Deal-413C56  no decision   buyer (back-to-school priority, CEO not ready)
Deal-47F1A1  competitor    buyer (WorkTango renewed 12 mo)
Deal-BF2A98  competitor    buyer (HiThrive deployed)
Deal-2A292B  other (build internally) buyer
Deal-D1AABF  no decision   unknown
Deal-FEDBCB  no decision   buyer (not engaged, reconnect EOY)
```

CATEGORY COUNTS (arithmetic)

```
no decision    33  (MIA/unresponsive 23 + pause/deprioritized/no-reason 10)
competitor     24
timing         18
pricing         5  (7ED004, 7B2236, C33D91, DAFB82, 8A119B)
product gap     4  (9048EB, 3618CC, 5AD03E, 981AD4)
other           5  (5DB9B0 spam, ED9AE7 multi-factor, 8E27DA scope, E0441F handoff, 2A292B build)
champion left   1  (F325A5)
Check: 33+24+18+5+4+5+1 = 90 ✓
```

SIDE SPLIT

```
buyer    60  = timing 18 + competitor 23 + no-decision 10 + pricing 5 + champion-left 1 + other 3
Bonusly   6  = product gap 4 + spam/not-ICP 1 + departed-rep handoff 1
unknown  24  = no-decision (pure MIA, no cause in text) 23 + competitor 1 (7CC678, no detail)
Check: 60+6+24 = 90 ✓
```

TAG vs FREE-TEXT CLEAR DISAGREEMENTS: 8

```
Deal-1BCA50  tag Competitor  · text "mostly about the budget"
Deal-9048EB  tag MIA         · text bad fit + multiple feature gaps
Deal-5E64CE  tag Doing nothing/Cost · text Nectar contract lock-in (competitor/incumbent)
Deal-3618CC  tag Lost DM     · text "Wanted Surveys" (product gap)
Deal-5AD03E  tag Competitor  · text "more defined budget access" (product gap, no vendor named)
Deal-8E27DA  tag Feature Request · text swag-provider-only scope decision, no feature request
Deal-ED9AE7  tag Lost DM     · text "Timing, budget, authority" (specific drivers)
Deal-55867E  tag Timing      · text generic decline, zero timing content
```
Borderline, not counted: Deal-381C8C / Deal-F1E8A6 / Deal-7CC678 (Competitor tags with silent text — absent evidence, not contradicting evidence).

TWO PATTERNS MOST WORTH ACTING ON

1. The engagement black hole — 33 no-decision deals (36.7% of losses), 23 with zero cause in the text ("unresponsive", "No contact after intro - ignored outreach from me and the ADR"). Side is unknown on 24 deals overall, meaning in over a quarter of losses the data cannot tell whether the buyer walked or Bonusly failed to engage. Several died at intro stage with 2-3 contacts — exploratory interest carried as pipeline. Act: qualification gate at intro (budget/timeline/initiative or disqualify), and a rule that "MIA" cannot be a closed-lost reason without stated cause — right now the largest bucket is unanalyzable by construction.

2. Incumbent inertia + regional fit in the competitor bucket — of 24 competitor losses, at least 7 are stay/renew-with-incumbent (Awardco A2C349, WorkTango 47F1A1, Kudos D0C698, Motivosity 64B19A, HiThrive BF2A98, Nectar 5E64CE, Paylocity 369281) and at least 4 are regional/currency fit (Rippl DDAB52 "without dealing with exchange rate differences", two Canadian providers 8A0992/9FCD0D, UK-focus gap in 981AD4). These are the two most repeatable competitive loss mechanisms in the file — displacement cost and non-US fit — and both are addressable with specific counter-positioning (switching-cost ROI case; international catalog/currency story) rather than generic differentiation.
pipeline-tieringPASS 1.00
Plain-English question: Score all 156 open deals into exactly one of LOCK / ACTION / BUILD / REVIVE / WATCH / RISKY.
Prompt — what the model saw
Score the open pipeline below on these signals: forecast category, stage, meetings_30d, contact count, and engagement recency. Assign exactly one tier per deal: LOCK, ACTION, BUILD, REVIVE, WATCH, or RISKY (RISKY = the forecast category disagrees with the engagement evidence). Never tier a deal LOCK if it has zero meetings_30d. Do NOT list every deal. Return ONLY a compact JSON object with exactly these keys: - "tier_counts": object mapping each tier to the number of deals in it (must sum to the total deal count), - "tier_examples": object mapping each tier to up to 3 example deal aliases from the data, - "risky_deals": array of the deal aliases you tiered RISKY, - "lock_violations": number of deals tiered LOCK that have zero meetings_30d (should be 0), - "pipeline_shape": one short paragraph on the overall pipeline shape. Note: inbound_emails_30d is 0 for every row due to a data defect, so treat meetings_30d as the inbound signal.
cost $1.1157 · 491s · in 84,434 / out 56,086 tok
✓ json_returned✓ tier_vocabulary✓ tier_counts_sum✓ risky_deals_listed✓ no_lock_violation_claimed

✓ Correct Answer (ground truth)

Plain-English answer: All 156 open deals tiered exactly once across the six tiers using forecast category, stage, meetings, contacts, and recency.
The answer in one line: total = 156
total156
graded checksjson_returned · tier_vocabulary · tier_counts_sum · risky_deals_listed · no_lock_violation_claimed

Model's actual answer

show the model's full answer
```json
{
  "tier_counts": {"LOCK": 5, "ACTION": 33, "BUILD": 68, "REVIVE": 3, "WATCH": 46, "RISKY": 1},
  "tier_examples": {
    "LOCK": ["Deal-D348E1", "Deal-C26D20", "Deal-403845"],
    "ACTION": ["Deal-25F752", "Deal-E53952", "Deal-5408B0"],
    "BUILD": ["Deal-D73B89", "Deal-012CB1", "Deal-523604"],
    "REVIVE": ["Deal-2D1F1B", "Deal-3EED2C", "Deal-57FF13"],
    "WATCH": ["Deal-6787C2", "Deal-0660B4", "Deal-BA571A"],
    "RISKY": ["Deal-7BBDFA"]
  },
  "risky_deals": ["Deal-7BBDFA"],
  "lock_violations": 0,
  "pipeline_shape": "Top-heavy and unqualified: 156 deals / $2.31M total, but 79% of value ($1.83M) sits in PIPELINE category with 71 of 156 deals at DS1–DS2, so the near-term quarter leans on a thin COMMIT base ($78K, 13 deals) plus $408K BEST_CASE. Tiering confirms the imbalance — only 5 LOCK (late-stage COMMIT with meeting evidence), while BUILD+WATCH hold 114 deals (73%) that have recency but insufficient meeting/velocity proof. Only 1 RISKY (Deal-7BBDFA: BEST_CASE $37,440, DS3, 0 meetings_30d, 0 touches in 30d, last contact 45 days stale — forecast category contradicted by engagement evidence). 3 REVIVE are cold DS1–DS2 PIPELINE deals with no touch since ~2026-06-16 or missing engagement rows. Meeting coverage is the core weakness: most DS3+ value advances on email-only cadence."
}
```

Data note: Deal-3EED2C and Deal-57FF13 have no row in engagements_by_deal_90d.csv — their engagement evidence is missing (only deals-file last_contacted_field used, both >30d stale → REVIVE). Tier counts verified: 5+33+68+3+46+1 = 156 = total deals. LOCK rule enforced: all 5 LOCK deals have meetings_30d ≥ 1 (Deal-D348E1: 1, Deal-C26D20: 4, Deal-403845: 2, Deal-A2B47C… excluded — zero-meeting COMMIT DS5 deals like Deal-547B2B fell to ACTION). inbound_emails_30d ignored per the stated defect; meetings_30d used as the inbound signal. Recency measured against the latest date in the data (2026-09-04).
call-transcript-extractionPASS 1.00
Plain-English question: Extract the CRM write-back fields from each call transcript as JSON.
Prompt — what the model saw
For each transcript, extract CRM write-back fields as JSON: why-buys (prospect statements only), pain points, stakeholders from the speaker list, budget signal (prospect-stated or null), timeline signal, competitor mentioned (only if the prospect raised it), next step (only if explicitly agreed), objections, confidence. Never fill a field from what the rep said or hoped.
cost $0.0645 · 55s · in 2,393 / out 3,765 tok
✓ json_emitted✓ budget_signal✓ null_discipline✓ stakeholders✓ competitor_prospect_only

✓ Correct Answer (ground truth)

Plain-English answer: Per transcript: prospect-only why-buys, pain points, stakeholders, budget and timeline signals, and competitors — e.g. TX-001: '$40k earmarked', live before open enrollment.
The answer in one line: rep = Alex Franklin
allowed_newTX-001, TX-002, TX-003, TX-004, TX-005, TX-006, Achievers, Bucketlist, Nectar
per_transcript{"TX-001": {"why_buys": ["automating anniversary and birthday awards"], "pain_points": ["HR team of three cannot keep up with it manually", "track everything in a spreadsheet, and people slip through
repAlex Franklin
rep_sourced_nulls{"TX-003": ["budget_signal \u2014 only the rep mentioned pricing/budget"], "TX-004": ["next_step \u2014 rep proposed, prospect said 'no promises'"], "TX-006": ["budget_signal \u2014 rep decoy, prospec
graded checksjson_emitted · budget_signal · null_discipline · stakeholders · competitor_prospect_only

Model's actual answer

show the model's full answer
```json
[
  {
    "transcript_id": "TX-001",
    "deal_alias": "Deal-CFE7F4",
    "why_buys": [
      "Automating anniversary and birthday awards — HR team of three cannot keep up with it manually (VP People)"
    ],
    "pain_points": [
      "Everything tracked in a spreadsheet; people slip through the cracks (HR Admin)",
      "Manual awards process exceeds HR team capacity (VP People)"
    ],
    "stakeholders": ["Prospect (VP People)", "Prospect (HR Admin)"],
    "budget_signal": "$40k earmarked for engagement tools this fiscal year (VP People)",
    "timeline_signal": "Live before open enrollment in November (VP People)",
    "competitor_mentioned": "Achievers — prospect-raised; evaluated last year, deemed too heavy for a team their size",
    "next_step": "Security review with IT on September 12 — explicitly agreed by VP People",
    "objections": [
      "Needs SSO and audit logs for IT sign-off (HR Admin)"
    ],
    "confidence": "high"
  },
  {
    "transcript_id": "TX-002",
    "deal_alias": "Deal-70BB30",
    "why_buys": [
      "Tie recognition to retention for hourly workforce — regretted turnover there is over 30% (Head of Total Rewards)"
    ],
    "pain_points": [
      "Regretted turnover over 30% among hourly workforce (Head of Total Rewards)"
    ],
    "stakeholders": ["Prospect (Head of Total Rewards)", "Prospect (CFO)"],
    "budget_signal": "$25k pilot budget approved by finance for this quarter (CFO)",
    "timeline_signal": "Decision by end of September (CFO)",
    "competitor_mentioned": null,
    "next_step": "Rep sends pilot agreement; CFO routes it to legal this week — explicitly agreed",
    "objections": [
      "Workday integration must be rock solid — CFO's stated one condition"
    ],
    "confidence": "high"
  },
  {
    "transcript_id": "TX-003",
    "deal_alias": "Deal-530B50",
    "why_buys": [
      "Make recognition visible across 12 retail locations (People Ops Manager)"
    ],
    "pain_points": [
      "Store managers have zero budget autonomy for on-the-spot recognition today (People Ops Manager)"
    ],
    "stakeholders": ["Prospect (People Ops Manager)"],
    "budget_signal": null,
    "timeline_signal": "No rush until Q1 (People Ops Manager)",
    "competitor_mentioned": "Bucketlist — prospect-raised; CEO used it at her last company and liked it",
    "next_step": "Schedule a call with the CEO; People Ops Manager will send two times — explicitly agreed",
    "objections": [
      "CEO must be sold first — she decides anything people-related and is not in the room (People Ops Manager)"
    ],
    "confidence": "medium"
  },
  {
    "transcript_id": "TX-004",
    "deal_alias": "Deal-180D02",
    "why_buys": [
      "Consolidate three separate recognition tools into one (VP People)"
    ],
    "pain_points": [
      "Paying for three tools and none of them talk to the HRIS (VP People)"
    ],
    "stakeholders": ["Prospect (VP People)", "Prospect (IT Security Lead)"],
    "budget_signal": "Under $15k annually = VP People can approve without going to the board (prospect-stated threshold, not a committed budget)",
    "timeline_signal": "Procurement cycle runs six to eight weeks minimum (IT Security Lead)",
    "competitor_mentioned": null,
    "next_step": null,
    "objections": [
      "Security review took three months for their last vendor — IT Security Lead's stated hesitation",
      "CFO follow-up not committed: 'Maybe — I need to check her calendar, no promises' (VP People)"
    ],
    "confidence": "medium"
  },
  {
    "transcript_id": "TX-005",
    "deal_alias": "Deal-F8767A",
    "why_buys": [
      "Automate service milestones (HR Director)",
      "Analytics on recognition equity across departments (HR Director)"
    ],
    "pain_points": [
      "Night-shift teams feel invisible — engagement scores run 20 points lower (People Ops Coordinator)"
    ],
    "stakeholders": ["Prospect (HR Director)", "Prospect (People Ops Coordinator)"],
    "budget_signal": "$12k approved under engagement line (HR Director)",
    "timeline_signal": "Running before the January all-hands (HR Director)",
    "competitor_mentioned": "Nectar — prospect-raised; mid-pilot with them right now",
    "next_step": "Rep presents directly to the exec team on October 2 — explicitly agreed by HR Director",
    "objections": [
      "Must beat the Nectar pilot experience (HR Director)",
      "Exec team skeptical after a failed rollout two years ago (HR Director)"
    ],
    "confidence": "medium"
  },
  {
    "transcript_id": "TX-006",
    "deal_alias": "Deal-EE195F",
    "why_buys": [
      "Cut admin time on service awards — personally spends five hours a month ordering and shipping plaques (HR Manager)"
    ],
    "pain_points": [
      "Five hours/month of manual plaque ordering and shipping (HR Manager)"
    ],
    "stakeholders": ["Prospect (HR Manager)"],
    "budget_signal": "No dollar figure stated; prospect said 'Budget isn't the issue — time is' (HR Manager)",
    "timeline_signal": "Fiscal year resets in February; a Q1 start is realistic (HR Manager)",
    "competitor_mentioned": null,
    "next_step": "Rep sends one-page overview; HR Manager forwards it to the COO this week — explicitly agreed",
    "objections": [
      "COO usually prefers building things in-house (HR Manager)",
      "Comparing against doing it internally — no other vendors in the mix (HR Manager)"
    ],
    "confidence": "medium"
  }
]
```

Extraction notes (arithmetic/decisions, data-bound):
- TX-003 budget_signal = null: the only price mentioned ($8/employee/month) was stated by the rep, not the prospect — excluded per instructions.
- TX-004 next_step = null: the CFO follow-up was explicitly non-committal ("Maybe… no promises"), so nothing was agreed.
- TX-002 and TX-006 competitor_mentioned = null: prospects stated no other vendors (TX-002: "first vendor we've had a real demo with"; TX-006: "Nobody else" — internal build is an alternative, not a named competitor).
- Stakeholder lists contain only speakers from the speaker list; the TX-003 CEO and TX-006 COO are referenced decision-makers but not speakers, so they appear only in objections/next steps where the prospect cited them.
- No dollar arithmetic was required beyond quoting stated figures verbatim ($40k, $25k, $15k threshold, $12k).
next-to-close0.75
Plain-English question: Which 3 deals are closest to signature right now?
Prompt — what the model saw
Which 3 deals are closest to signature right now? Use the CRM data and the Slack thread; Slack often has the real status before the CRM does. For each: deal, amount, why it is close, what is left. Under 250 words.
cost $0.0629 · 33s · in 4,969 / out 1,894 tok
✓ exactly_three✗ top3_named✓ slack_first✓ blocked_excluded

✓ Correct Answer (ground truth)

Plain-English answer: Deal-547B2B ($11,200), Deal-B7EBD1 ($9,000), and Deal-403845 ($9,000) — all DS5 COMMIT deals closing September 10-11, with what is left on each.
blocked_excludeddeal_alias=Deal-2465CE; reason=champion left + procurement freeze (Slack #sales-sparring 2026-09-04)
exactly_three1
ranking_ruleSlack signature-imminent deal ranks first; remaining ranked by stage (DS5>DS4), forecast (COMMIT>BEST_CASE>PIPELINE), close_date asc, amount desc
top3rank=1; deal_alias=Deal-547B2B; amount=11200; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11 | rank=2; deal_alias=Deal-B7EBD1; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-10 | rank=3; deal_alias=Deal-403845; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11
graded checksexactly_three · top3_named · slack_first · blocked_excluded

Model's actual answer

show the model's full answer
Top 3 closest to signature (CRM + Slack, Slack overrides CRM per freshness):

1. Deal-547B2B · $11,200 · DS5/COMMIT · Alex Franklin
Slack 09-04: redlines came back clean, signing page is out, VP People confirmed signing "tomorrow." Why close: legal done, document sent, buyer named a signature date. Left: capture the executed signature and confirm it actually landed, since no later data confirms closure.

2. Deal-403845 · $9,000 · DS5/COMMIT · Alex Franklin (owner discrepancy: Dana Mercer posted it as "my side" on 09-04, CRM lists Franklin; no data resolves this)
Why close: Slack 09-04 says it is "moving fine" and the order form is already with the customer's finance team. Left: finance sign-off and returned order form.

3. Deal-A2B47C · $6,360 · DS5/COMMIT · Alex Franklin
Why close: Slack 09-04 confirms "still warm, just normal legal-review pace," no blockers. Left: legal review completion, then signature.

Excluded: Deal-2465CE ($5,400, DS5/COMMIT) despite CRM status. Slack 09-04: champion left, procurement froze new vendors, rep pulling it from commit, now a Q4 deal. Deal-D348E1 ($13,770) is warm per Slack but closes 10-15, later than A2B47C's 09-11.

Arithmetic: ranked amount ordering among the three is $11,200 > $9,000 > $6,360; total = $26,560. Ranking was driven by Slack signal strength, not amount.

Data gaps, stated explicitly: no engagement, reply, meeting, or deal-room data provided; CRM close dates (09-10, 09-11) predate the Slack export's implied follow-up, and nothing in either file confirms whether these signed after 09-04.

Next action: confirm signature receipt on Deal-547B2B with Alex Franklin today.
product-gap-detectionPASS 1.00
Plain-English question: Which prospect complaints are true product gaps versus rollout or enablement issues?
Prompt — what the model saw
Surface true product gaps raised by prospects in active deals. For each candidate: quote the prospect line, classify as real gap, plan gate, or rollout/enablement issue using the product docs excerpt, and name the deal and amount. Only real gaps go in the summary. Rep language does not count as prospect voice.
cost $0.0284 · 19s · in 1,089 / out 1,191 tok
✓ real_gaps✓ plan_gate_classified✓ enablement_classified✓ gap_summary

✓ Correct Answer (ground truth)

Plain-English answer: Two real gaps (SCIM provisioning, ADP Workforce Now integration) separated from plan gates and enablement issues, each with the prospect quote and deal.
allowed_newTG-001, TG-002, TG-003, TG-004, TG-005, SCIM user provisioning, ADP Workforce Now integration, Custom report builder
enablement_issuestranscript_id=TG-004; deal_alias=Deal-84DBA6; amount=16000; feature=Slack integration recognition adoption; quote=We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.; why_enablement=Slack integration exists on all plans; managers were never trained
excluded_rep_voicetranscript_id=TG-005; line=Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap.; reason=rep voice, not prospect-raised
plan_gatestranscript_id=TG-003; deal_alias=Deal-CFE7F4; amount=32000; feature=Custom report builder; tier_required=Enterprise; quote=I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?
real_gapstranscript_id=TG-001; deal_alias=Deal-EC3025; amount=62000; feature=SCIM user provisioning; quote=We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.; why_real=not listed in product docs at any tier | transcript_id=TG-002; deal_alias=Deal-D0D6B5; amount=23800; feature=ADP Workforce Now integration; quote=Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.; why_real=supported HRIS list is Workday/BambooHR/Gusto/Rippling only
graded checksreal_gaps · plan_gate_classified · enablement_classified · gap_summary

Model's actual answer

show the model's full answer
Product gap triage — transcripts_gaps.csv (5 transcripts) vs product_docs.md (snapshot 2026-08)

Missing data notice: the provided files contain no deal amounts. I cannot name amounts without inventing them, so every deal below is cited by alias only.

---

TG-001 — Deal-EC3025 — REAL GAP
Prospect (IT Security Lead): "We need SCIM user provisioning — our security team will not approve any tool that requires manual account management."
Docs check: product_docs.md states "SCIM user provisioning ... NOT currently listed as supported capabilities." Not offered on any tier (Core/Pro/Enterprise).
Classification: Real gap — capability does not exist at any plan level.
Amount: not provided in data.

TG-002 — Deal-D0D6B5 — REAL GAP
Prospect (HRIS Manager): "Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us."
Docs check: Supported HRIS integrations are Workday, BambooHR, Gusto, Rippling (Pro+). Docs state "ADP Workforce Now integrations are NOT currently listed as supported capabilities."
Classification: Real gap — ADP is absent from the supported integration list entirely.
Amount: not provided in data.

TG-003 — Deal-CFE7F4 — PLAN GATE (not a gap)
Prospect (VP People): "I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?"
Docs check: Custom report builder = Enterprise tier only ("— / — / yes" in the plan matrix). The capability exists; it's gated by plan.
Classification: Plan gate.
Amount: not provided in data.

TG-004 — Deal-84DBA6 — ROLLOUT/ENABLEMENT (not a gap)
Prospect (People Ops Manager): "We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it."
Docs check: Slack integration is supported on all plans ("Slack, Microsoft Teams (all plans)"). The product works as documented; adoption failed due to missing manager training.
Classification: Rollout/enablement issue.
Amount: not provided in data.

TG-005 — Deal-36C33F — EXCLUDED (rep language, not prospect voice)
The mobile-app admission ("we don't have a native mobile app for hourly workers yet, though it's on the roadmap") comes from Alex Franklin, the rep. The prospect's only line is: "Good to know. The web version should be fine for our office staff for now." No gap raised by the prospect. Per the rules, rep language does not count.

---

SUMMARY — true product gaps only

1. Deal-EC3025 — SCIM user provisioning missing at all tiers. Blocking IT security approval. Amount: not in provided data.
2. Deal-D0D6B5 — ADP Workforce Now integration missing. Prospect explicitly called it a dealbreaker. Amount: not in provided data.

Counts: 5 transcripts reviewed → 2 real gaps, 1 plan gate, 1 rollout/enablement, 1 excluded (rep-voiced).
rep-performance 5 tests
stale-pipeline-by-rep0.83
Plain-English question: Which open deals have had no email, call, or meeting in the last 7 days?
Prompt — what the model saw
The snapshot date is 2026-09-05. Using the open deals and the engagements-by-deal table, list every open deal with no logged email, call, or meeting in the last 7 days. Compute recency from the engagements table's last_email/last_call/last_meeting fields, NOT from any date field on the deal (the deal's last_contacted_field is unreliable). For each stale deal give: deal alias, owner name, stage, amount, days since last contact. Group by owner, order by amount descending within each owner. Then state, per owner, how many stale deals they carry and the total stale amount. (Do not write outreach text.)
cost $0.6417 · 299s · in 56,973 / out 31,796 tok
✓ stale_recall_ge_80✓ grouped_by_owner✓ jake_top_amount✓ kevin_stale_amt✓ jenn_stale_amt✗ stale_count

✓ Correct Answer (ground truth)

Plain-English answer: 65 open deals are stale; the largest is Deal-2D1F1B at $240,000, with Bryce Harmon carrying the most stale dollars at $626,243.
The answer in one line: jake_amt = 626,243; jenn_amt = 261,645; kevin_amt = 252,905.03; top_amount = 240,000; stale_count = 65; top_alias = Deal-2D1F1B
jake_amt626,243
jenn_amt261,645
kevin_amt252,905.03
stale_count65
top_amount240,000
top_aliasDeal-2D1F1B
graded checksstale_recall_ge_80 · grouped_by_owner · jake_top_amount · kevin_stale_amt · jenn_stale_amt · stale_count

Model's actual answer

show the model's full answer
Stale-deal audit, snapshot 2026-09-05. Method: for each open deal, last contact = max(last_email, last_call, last_meeting) from engagements_by_deal_90d.csv; days since = 2026-09-05 minus that date; stale = days > 7 (i.e., nothing on/after 2026-08-29). The deal's last_contacted_field was ignored as instructed. 156 open deals scanned; 63 stale.

DATA GAPS (cannot be evaluated — no row in the engagements table):
- Deal-3EED2C (Alex Franklin, DS2, $7,200) — no engagement row exists; last_contacted_field on the deal is blank too.
- Deal-57FF13 (Elena Sinclair, DS1, $2,100) — no engagement row exists.
These are excluded from the counts below and flagged as unknown, not fresh.

== Bryce Harmon (owner_id 119337721) ==
| Deal | Stage | Amount | Days since last contact | Last contact |
|---|---|---|---|---|
| Deal-2D1F1B | DS1 | $240,000.00 | 81 | 2026-06-16 (meeting) |
| Deal-66D1FC | DS1 | $99,000.00 | 16 | 2026-08-20 (email) |
| Deal-950043 | DS1 | $70,000.00 | 19 | 2026-08-17 (email) |
| Deal-B23205 | DS1 | $45,000.00 | 16 | 2026-08-20 (email) |
| Deal-7BBDFA | DS3 | $37,440.00 | 46 | 2026-07-21 (email) |
| Deal-332637 | DS2 | $36,000.00 | 9 | 2026-08-27 (email) |
| Deal-1BEEBF | DS1 | $31,500.00 | 19 | 2026-08-17 (email) |
| Deal-C5658B | DS1 | $23,400.00 | 16 | 2026-08-20 (email) |
| Deal-40522D | DS3 | $21,000.00 | 19 | 2026-08-17 (email) |
| Deal-F0EBBB | DS3 | $11,400.00 | 24 | 2026-08-12 (email) |
| Deal-E25A09 | DS1 | $6,000.00 | 9 | 2026-08-27 (email) |
| Deal-C9C286 | DS2 | $5,502.00 | 9 | 2026-08-27 (email) |
| Deal-012CB1 | DS1 | $1.00 | 23 | 2026-08-13 (email) |
Stale count: 13. Stale amount: 240,000 + 99,000 + 70,000 + 45,000 + 37,440 + 36,000 + 31,500 + 23,400 + 21,000 + 11,400 + 6,000 + 5,502 + 1 = $626,243.00

== Dana Mercer (owner_id 83155923) ==
| Deal | Stage | Amount | Days | Last contact |
|---|---|---|---|---|
| Deal-44EA29 | DS2 | $60,000.00 | 10 | 2026-08-26 (email) |
| Deal-E51FB7 | DS2 | $43,875.00 | 12 | 2026-08-24 (call) |
| Deal-B42F46 | DS1 | $27,000.00 | 19 | 2026-08-17 (email) |
| Deal-BA3DDC | DS3 | $23,400.00 | 15 | 2026-08-21 (call) |
| Deal-9DDE86 | DS2 | $20,000.00 | 15 | 2026-08-21 (email) |
| Deal-215CCA | DS3 | $18,900.00 | 17 | 2026-08-19 (meeting) |
| Deal-5EED42 | DS3 | $16,250.00 | 11 | 2026-08-25 (email/call) |
| Deal-57887A | DS2 | $15,000.00 | 8 | 2026-08-28 (email) |
| Deal-B7EBD1 | DS5 | $9,000.00 | 16 | 2026-08-20 (email) |
| Deal-3974EB | DS4 | $9,000.00 | 8 | 2026-08-28 (email/meeting) |
| Deal-F40F04 | DS2 | $8,100.00 | 15 | 2026-08-21 (email/meeting) |
| Deal-87DDD1 | DS1 | $5,000.00 | 19 | 2026-08-17 (email) |
| Deal-F336B6 | DS3 | $4,200.00 | 15 | 2026-08-21 (email) |
| Deal-0660B4 | DS4 | $1,920.00 | 16 | 2026-08-20 (meeting) |
Stale count: 14. Stale amount: 60,000 + 43,875 + 27,000 + 23,400 + 20,000 + 18,900 + 16,250 + 15,000 + 9,000 + 9,000 + 8,100 + 5,000 + 4,200 + 1,920 = $261,645.00

== Alex Franklin (owner_id 84342457) ==
| Deal | Stage | Amount | Days | Last contact |
|---|---|---|---|---|
| Deal-CC08D1 | DS1 | $24,000.00 | 16 | 2026-08-20 (email) |
| Deal-E73427 | DS3 | $18,000.00 | 10 | 2026-08-26 (email/meeting) |
| Deal-885F45 | DS2 | $9,300.00 | 12 | 2026-08-24 (email) |
| Deal-C2FF3C | DS1 | $8,316.00 | 10 | 2026-08-26 (email) |
| Deal-0D2F7A | DS3 | $5,100.00 | 12 | 2026-08-24 (call) |
| Deal-6C60D4 | DS3 | $4,800.00 | 12 | 2026-08-24 (call) |
| Deal-13FEBD | DS2 | $4,680.00 | 12 | 2026-08-24 (call) |
| Deal-9D0060 | DS3 | $3,840.00 | 12 | 2026-08-24 (email) |
| Deal-690476 | DS2 | $3,600.00 | 18 | 2026-08-18 (call) |
| Deal-C6D97A | DS4 | $3,240.00 | 8 | 2026-08-28 (email) |
| Deal-EE195F | DS3 | $3,120.00 | 8 | 2026-08-28 (email) |
| Deal-278DEC | DS3 | $2,700.00 | 8 | 2026-08-28 (email) |
| Deal-635B8E | DS3 | $2,600.00 | 18 | 2026-08-18 (email) |
| Deal-6883F3 | DS1 | $2,400.00 | 16 | 2026-08-20 (email/meeting) |
| Deal-4A13AD | DS3 | $2,160.00 | 26 | 2026-08-10 (email) |
| Deal-F67D31 | DS2 | $1,800.00 | 8 | 2026-08-28 (email) |
| Deal-5FDCE4 | DS3 | $1,600.00 | 12 | 2026-08-24 (email) |
| Deal-BA571A | DS4 | $1,080.00 | 18 | 2026-08-18 (email) |
Stale count: 18. Stale amount: 24,000 + 18,000 + 9,300 + 8,316 + 5,100 + 4,800 + 4,680 + 3,840 + 3,600 + 3,240 + 3,120 + 2,700 + 2,600 + 2,400 + 2,160 + 1,800 + 1,600 + 1,080 = $102,336.00
(Plus Deal-3EED2C, $7,200 — unevaluable, see data gaps.)

== Cole Ingram (owner_id 83155924) ==
| Deal | Stage | Amount | Days | Last contact |
|---|---|---|---|---|
| Deal-D04904 | DS2 | $58,529.25 | 11 | 2026-08-25 (email) |
| Deal-B25F40 | DS3 | $40,000.00 | 8 | 2026-08-28 (email) |
| Deal-813836 | DS2 | $32,175.00 | 11 | 2026-08-25 (email) |
| Deal-1BA595 | DS2 | $31,750.00 | 11 | 2026-08-25 (email) |
| Deal-CFE1E8 | DS3 | $18,000.00 | 11 | 2026-08-25 (email) |
| Deal-CD47A6 | DS2 | $12,168.00 | 11 | 2026-08-25 (email) |
| Deal-627646 | DS3 | $11,193.00 | 11 | 2026-08-25 (email) |
| Deal-FF809F | DS2 | $7,781.20 | 11 | 2026-08-25 (email) |
| Deal-AF932D | DS2 | $7,225.40 | 11 | 2026-08-25 (email) |
| Deal-A71728 | DS2 | $6,947.50 | 11 | 2026-08-25 (email) |
| Deal-8BC9F5 | DS2 | $5,616.00 | 10 | 2026-08-26 (email) |
| Deal-175395 | DS3 | $4,779.88 | 11 | 2026-08-25 (email) |
| Deal-481E24 | DS3 | $4,140.00 | 10 | 2026-08-26 (call) |
| Deal-C7F9BF | DS2 | $3,360.00 | 11 | 2026-08-25 (email) |
| Deal-2F3A66 | DS3 | $3,334.80 | 11 | 2026-08-25 (email) |
| Deal-342E96 | DS2 | $2,700.00 | 24 | 2026-08-12 (email) |
| Deal-E568D5 | DS3 | $1,875.00 | 11 | 2026-08-25 (email) |
| Deal-FD9F4E | DS5 | $1,330.00 | 10 | 2026-08-26 (email) |
Stale count: 18. Stale amount: 58,529.25 + 40,000 + 32,175 + 31,750 + 18,000 + 12,168 + 11,193 + 7,781.20 + 7,225.40 + 6,947.50 + 5,616 + 4,779.88 + 4,140 + 3,360 + 3,334.80 + 2,700 + 1,875 + 1,330 = $252,905.03

== Farid Osman (owner_id 716654662) ==
No stale deals. All 6 of his open deals show email activity within the window (last_email 2026-09-01 through 2026-09-04, all ≤ 4 days). Stale count: 0. Stale amount: $0.

== Elena Sinclair (owner_id 701163055) ==
Her only open deal (Deal-57FF13, DS1, $2,100) has no engagement row — status unknown, cannot confirm stale or fresh. Stale count: 0 confirmed (1 unevaluable). Stale amount: $0 confirmed.

Per-owner summary:
| Owner | Stale deals | Stale amount |
|---|---|---|
| Bryce Harmon | 13 | $626,243.00 |
| Dana Mercer | 14 | $261,645.00 |
| Alex Franklin | 18 (+1 unevaluable) | $102,336.00 |
| Cole Ingram | 18 | $252,905.03 |
| Farid Osman | 0 | $0 |
| Elena Sinclair | 0 confirmed (1 unevaluable) | $0 confirmed |
| TOTAL | 63 | $1,243,129.03 |

Notes on the arithmetic: days = (2026-09-05 − last contact date), e.g., Deal-2D1F1B: 2026-06-16 → 81 days; Deal-57887A: 2026-08-28 → 8 days. Several engagement rows contain future-dated meetings (e.g., Deal-944310 last_meeting 2026-09-15); those deals counted as fresh since a contact exists on record, but the future dates look like data-entry artifacts worth flagging.
activity-mix-vs-outcomePASS 1.00
Plain-English question: Per rep, what did the last 30 days of emails, calls, and meetings look like next to the deals that advanced?
Prompt — what the model saw
The snapshot date is 2026-09-05. For each sales rep (owner), compute the last-30-day totals of emails, calls, and meetings from the engagements table, and the number of deals that entered DS2 in the last 30 days (t_ds2 within 30 days of the snapshot). Report per rep: the activity mix as percentages (emails/calls/meetings share of that rep's total activities), total activities per DS2 entry, and rank the reps by that efficiency ratio (lowest activities per DS2 entry = most efficient). Name the most efficient rep and the highest-volume rep, and note if they differ. Do not attribute cause.
cost $0.5967 · 272s · in 55,229 / out 28,631 tok
✓ alex_ds2_30d✓ justin_volume_high✓ efficient_named✓ mix_present✓ no_causal_claim

✓ Correct Answer (ground truth)

Plain-English answer: Alex Franklin is the standout signal with 18 deals entering DS2 in 30 days (next closest: Bryce Harmon with 4), alongside per-rep activity totals.
The answer in one line: alex_ds2_30d = 18; bryce_ds2_30d = 4; note = legacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
alex_ds2_30d18
bryce_ds2_30d4
notelegacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
graded checksalex_ds2_30d · justin_volume_high · efficient_named · mix_present · no_causal_claim

Model's actual answer

show the model's full answer
SNAPSHOT 2026-09-05 · 30-day window = 2026-08-06 through 2026-09-05 (t_ds2 inclusive)

DATA CAVEATS
- Deal-3EED2C (Alex Franklin) and Deal-57FF13 (Elena Sinclair) have no row in engagements_by_deal_90d.csv. Their activity counts as 0. Deal-3EED2C's t_ds2=2026-09-03 IS counted as a DS2 entry.
- Elena Sinclair: her only deal (Deal-57FF13) has no engagement row and no t_ds2 → 0 activities, 0 DS2 entries, ratio undefined. Excluded from ranking.

PER REP (activities summed from emails_30d + calls_30d + meetings_30d across that rep's deals)

1) Alex Franklin (84342457)
   Emails 307 + Calls 36 + Meetings 41 = 384 total
   Mix: 307/384=79.9% · 36/384=9.4% · 41/384=10.7%
   DS2 entries in window: 18 (Deal-EE195F, Deal-D9A72E 08-06; Deal-7FA0C3, Deal-E531A6 08-07; Deal-36C33F, Deal-D1E6C2 08-11; Deal-317E6F 08-12; Deal-4F775F 08-17; Deal-F436DA 08-19; Deal-CA5E44 08-24; Deal-46988D 08-26; Deal-5296C9, Deal-898FC5, Deal-E73427 08-28; Deal-403845, Deal-92D97D 09-02; Deal-1FC049, Deal-3EED2C 09-03)
   Ratio: 384/18 = 21.33 activities per DS2 entry

2) Bryce Harmon (119337721)
   Emails 162 + Calls 0 + Meetings 43 = 205 total
   Mix: 162/205=79.0% · 0.0% · 43/205=21.0%
   DS2 entries: 4 (Deal-1CCE5C 08-06, Deal-25F752 08-10, Deal-CA7DC0 08-12, Deal-D73B89 09-03)
   Ratio: 205/4 = 51.25

3) Dana Mercer (83155923)
   Emails 84 + Calls 18 + Meetings 11 = 113 total
   Mix: 84/113=74.3% · 18/113=15.9% · 11/113=9.7%
   DS2 entries: 1 (Deal-57887A 08-07)
   Ratio: 113/1 = 113.00

4) Cole Ingram (83155924)
   Emails 96 + Calls 14 + Meetings 1 = 111 total
   Mix: 96/111=86.5% · 14/111=12.6% · 1/111=0.9%
   DS2 entries: 2 (Deal-1BA595 08-12, Deal-42326B 08-26)
   Ratio: 111/2 = 55.50

5) Farid Osman (716654662)
   Emails 38 + Calls 0 + Meetings 34 = 72 total
   Mix: 38/72=52.8% · 0.0% · 34/72=47.2%
   DS2 entries: 1 (Deal-499BF6 08-26)
   Ratio: 72/1 = 72.00

6) Elena Sinclair (701163055)
   No engagement data, no DS2 entries → ratio undefined.

EFFICIENCY RANKING (lowest activities per DS2 entry = most efficient)
   1. Alex Franklin — 21.33
   2. Bryce Harmon — 51.25
   3. Cole Ingram — 55.50
   4. Farid Osman — 72.00
   5. Dana Mercer — 113.00

ANSWER
- Most efficient rep: Alex Franklin (21.33 activities per DS2 entry)
- Highest-volume rep: Alex Franklin (384 activities in 30 days)
- They do NOT differ — the same rep leads both.
ae-qtd-scorecardPASS 1.00
Plain-English question: Build Alex Franklin's QTD scorecard: bookings versus quota, pipeline, and conversion.
Prompt — what the model saw
Build a QTD scorecard for the named AE as of 2026-09-05: bookings vs quota with attainment percent, new vs expansion split, active pipeline by stage with amount, rolling 90-day DS2-to-won rate, win and loss counts with top loss reason, activity volume by type in the last 30 days. Exclude any closed-won deal dated before the quarter. Close with three coaching observations grounded in the numbers.
cost $0.5900 · 237s · in 35,271 / out 22,570 tok
✓ bookings✓ quota✓ attainment✓ new_expansion_split✓ top_loss_reason✓ coaching_obs

✓ Correct Answer (ground truth)

Plain-English answer: 75% attainment ($150,000 on a $200,000 quota), $113,500 new versus $36,500 expansion, pipeline by stage, and the rolling DS2-to-won rate.
The answer in one line: quota = 200000; bookings = 150000; new_bookings = 113500; expansion_bookings = 36500; open_deals_count = 125; attainment_pct = 75; owner = Alex Franklin (+12 more below)
quota200,000
attainment_pct75
bookings150,000
expansion_bookings36,500
expansion_count3
losses_qtd27
new_bookings113,500
new_count5
open_deals_count125
wins_qtd8
activity_30d_totalsemails_30d=807; calls_30d=112; meetings_30d=128; notes_30d=50
allowed_new18 items — e.g. Deal-A1C3E5, Deal-B7D2F4, Deal-C9E1A6, Deal-D4B8C2
ds2_to_wonwindow_start=2026-06-07; won_in_window=8; closed_in_window=35; rate=0.2286
ownerAlex Franklin
pipeline_by_stage{"DS1": {"count": 20, "amount": 284621}, "DS2": {"count": 28, "amount": 353760}, "DS3": {"count": 67, "amount": 552705}, "DS4": {"count": 5, "amount": 23574}, "DS5": {"count": 5, "amount": 45730}}
prior_quarter_excludeddeal_alias=Deal-B3E6F1; close_date=2026-06-20; amount=24000
quarter2026-Q3
snapshot_date2026-09-05
top_loss_reasonreason=Lost- Timing (1 year or more); count=13
graded checksbookings · quota · attainment · new_expansion_split · top_loss_reason · coaching_obs

Model's actual answer

show the model's full answer
QTD SCORECARD — Alex Franklin — as of 2026-09-05 (Q3 2026: quarter start 2026-07-01)
Sources: ae_deals.csv, ae_engagements.csv, quota.csv. All figures computed from those files only.

1) BOOKINGS VS QUOTA
Quota (quota.csv, 2026-Q3, snapshot 2026-09-05): $200,000
QTD closed-won (close_date 2026-07-01 through 2026-09-05; Deal-B3E6F1, $24,000, closed 2026-06-20, EXCLUDED as pre-quarter):
- Deal-A1C3E5 $40,000 (2026-07-15, new)
- Deal-F2C7D8 $20,000 (2026-07-24, expansion)
- Deal-B7D2F4 $35,000 (2026-07-31, new)
- Deal-C9E1A6 $21,000 (2026-08-12, new)
- Deal-A8B4D6 $12,000 (2026-08-19, expansion)
- Deal-D4B8C2 $11,000 (2026-08-21, new)
- Deal-E6F3A9 $6,500 (2026-09-02, new)
- Deal-C5D9E2 $4,500 (2026-09-03, expansion)
Bookings = 40,000+20,000+35,000+21,000+12,000+11,000+6,500+4,500 = $150,000 (8 deals)
Attainment = 150,000 / 200,000 = 75.0%

2) NEW VS EXPANSION SPLIT (QTD won)
New: A1C3E5 40,000 + B7D2F4 35,000 + C9E1A6 21,000 + D4B8C2 11,000 + E6F3A9 6,500 = $113,500 (5 deals, 75.7%)
Expansion: F2C7D8 20,000 + A8B4D6 12,000 + C5D9E2 4,500 = $36,500 (3 deals, 24.3%)

3) ACTIVE PIPELINE BY STAGE (125 open deals, $1,260,390 total)
DS1: 20 deals, $284,621
DS2: 28 deals, $353,760
DS3: 67 deals, $552,705
DS4: 5 deals, $23,574
DS5: 5 deals, $45,730
Check: 284,621+353,760+552,705+23,574+45,730 = $1,260,390

4) ROLLING 90-DAY DS2-TO-WON RATE (entered_ds2 between 2026-06-07 and 2026-09-05)
All 8 QTD won deals entered DS2 in-window (dates 2026-06-22 through 2026-08-10). All 27 closed-lost deals entered DS2 in-window (2026-06-12 through 2026-08-08). No won or lost deal falls outside the window. 76 open deals also entered DS2 in-window (unresolved, excluded from a closed-rate calc).
Closed-rate: 8 won / (8 won + 27 lost) = 8/35 = 22.9%
Including still-open entrants: 8/111 = 7.2% (informational; unresolved)

5) WINS AND LOSSES (QTD, closed in quarter through 2026-09-05)
Wins: 8 ($150,000)
Losses: 27 ($344,852 in lost amount; all 27 loss close dates 2026-07-29 through 2026-09-02)
Loss reasons: Lost- Timing (1 year or more) 13; MIA 5; Competitor 5; Lost DM 2; Feature Request 1 (Deal-8E27DA); Lost- Does not fit ICP 1 (Deal-5DB9B0)
Top loss reason: "Lost- Timing (1 year or more)" — 13/27 = 48.1%
Win rate on resolved deals: 8/35 = 22.9%

6) ACTIVITY VOLUME, LAST 30 DAYS (ae_engagements.csv, summed across all 161 deals)
Emails: 807
Calls: 112
Meetings: 128
Notes: 50
Total: 807+112+128+50 = 1,097 touches
Segment view: 9 won deals = 181 touches (avg 20.1/deal); 27 lost deals = 172 (avg 6.4); 125 open deals = 744 (avg 6.0).

COACHING OBSERVATIONS
1) Qualification is the leak, not effort: 48.1% of losses (13/27) cite "Lost- Timing (1 year or more)" and the closed DS2-to-won rate is only 22.9% (8/35). Deals are entering DS2 that were never going to buy this cycle — tighten DS2 entry criteria (budget window, decision timeline) so pipeline reflects real quarters.
2) Engagement depth separates outcomes: won deals averaged 20.1 touches in 30 days (incl. 33 calls, 25 meetings across 9 won deals) vs 6.4 for lost deals (13 meetings across 27). Yet the 125 open deals average just 6.0 touches and only 1 note total — the open book is being worked at lost-deal intensity, and next-step documentation is essentially absent.
3) Attainment math is tight against a back-loaded pipe: at 75.0% with ~$50,000 needed in the final 25 days, near-term cover is thin — DS4+DS5 hold only $69,304 combined (10 deals), and $1,191,086 (94.5%) of open pipeline sits in DS1–DS3 with close dates mostly in October–December. September push should concentrate on the DS5 deals (Deal-403845 $9,000, Deal-547B2B $11,200, Deal-A2B47C $6,360, all closing 2026-09-11) and DS4 (Deal-1FC049 $1,920, Deal-F9A08A $2,484), not the DS1–DS3 bulk.
multithreading-gapPASS 1.00
Plain-English question: Which open deals are single-threaded or under-threaded?
Prompt — what the model saw
Find every open deal that is single-threaded (fewer than 2 active contacts) or under-threaded (fewer than 3, or all contacts in one persona). Active means engaged in the last 60 days and not marked former. Personas: economic buyer, champion, HR admin, IT security, finance. For each flagged deal: amount, stage, active contact count, personas present, personas missing, the most valuable persona to add given the stage, and any on-file unengaged contact at that company who fits; otherwise say none on file.
cost $0.0670 · 56s · in 2,633 / out 3,896 tok
✓ flagged_recall✓ persona_language✓ single_thread_term✓ on_file_fit

✓ Correct Answer (ground truth)

Plain-English answer: Deals with fewer than 2 active contacts (single-threaded) or fewer than 3, or all in one persona (under-threaded), each flagged with the missing persona.
The answer in one line: active_window_start = 2026-07-07
active_window_start2026-07-07
allowed_new59 items — e.g. C-077A0E, C-0D15DF, C-1BDFE0, C-2AE3AA
best_persona_ruleDS2/DS3: economic buyer > IT security > finance > HR admin > champion; DS4/DS5: economic buyer > finance > IT security > HR admin > champion; first persona not present among active contacts
flag_ruleflag if active_contacts < 2 (single-threaded) or active_contacts < 3 or all active contacts share one persona (under-threaded); active = engaged within 60 days and not former
flaggeddeal_alias=Deal-36C33F; amount=15000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-50D386; amount=36000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5408B0; amount=14850; stage=DS4; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5BFE3B; amount=18000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-885F45; amount=9300; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=IT security | deal_alias=Deal-92D97D; amount=60000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer … (+5 more)
healthy_not_flaggedDeal-4B0BEB, Deal-84DBA6, Deal-D348E1
persona_universeeconomic buyer, champion, HR admin, IT security, finance
graded checksflagged_recall · persona_language · single_thread_term · on_file_fit

Model's actual answer

show the model's full answer
THREADING AUDIT — open deals, single/under-threaded

METHOD (arithmetic shown)
- Today = 2026-09-24. Active window = last 60 days → cutoff 2026-09-24 − 60 days = 2026-07-26. A contact is ACTIVE if last_engaged_date ≥ 2026-07-26 AND is_former = false.
- Flag rules: single-threaded = <2 active contacts; under-threaded = <3 active contacts OR all active contacts share one persona.
- DATA MISSING: The files contain no deal amount and no stage field, and no open/closed indicator. Amount = NOT PROVIDED; Stage = NOT PROVIDED for every deal below. I treat all listed deals as in-scope (nothing marks them closed). Because stage is absent, "most valuable persona to add" is reasoned from persona coverage, not stage.

Deals audited: 14. Flagged: 11. Clean: 3 (Deal-84DBA6, Deal-4B0BEB, Deal-D348E1).

FLAGGED DEALS

1. Deal-EC3025 (C-FDD0C7) — single-threaded
   Amount: not provided | Stage: not provided
   Active: 1 of 2 (CT-F2C1AE, economic buyer, is_former=true → excluded)
   Personas present: champion | Missing: economic buyer, HR admin, IT security, finance
   Add: economic buyer (no active/future buyer at all)
   On file: CT-6827DB, Chief People Officer, economic buyer

2. Deal-92D97D (C-E23238) — single-threaded
   Amount: not provided | Stage: not provided
   Active: 1 of 2 (CT-A902AE champion last engaged 2026-06-01 < 2026-07-26 → stale, excluded; not former, so re-engageable)
   Personas present: HR admin | Missing: economic buyer, champion, IT security, finance
   Add: economic buyer
   On file: none on file

3. Deal-50D386 (C-EB10E4) — under-threaded (2 active < 3)
   Amount: not provided | Stage: not provided
   Active: 2 of 2 (champion 2026-09-01; HR admin 2026-08-25)
   Personas present: champion, HR admin | Missing: economic buyer, IT security, finance
   Add: economic buyer
   On file: CT-A1C4B3, Chief People Officer, economic buyer

4. Deal-D0D6B5 (C-32918E) — under-threaded (all 3 active = one persona)
   Amount: not provided | Stage: not provided
   Active: 3 of 3 (champion ×3: 2026-09-02, 2026-08-19, 2026-08-07)
   Personas present: champion | Missing: economic buyer, HR admin, IT security, finance
   Add: economic buyer
   On file: CT-1FA4DB, Chief People Officer, economic buyer

5. Deal-5BFE3B (C-535D36) — under-threaded (2 active < 3 AND single persona)
   Amount: not provided | Stage: not provided
   Active: 2 of 2 (champion ×2: 2026-08-31, 2026-08-12)
   Personas present: champion | Missing: economic buyer, HR admin, IT security, finance
   Add: economic buyer
   On file: none on file

6. Deal-36C33F (C-077A0E) — single-threaded
   Amount: not provided | Stage: not provided
   Active: 1 of 3 (IT security 2026-08-15; CT-405B45 champion and CT-86B22F economic buyer both is_former=true → excluded)
   Personas present: IT security | Missing: economic buyer, champion, HR admin, finance
   Add: economic buyer (both prior buyer and champion departed; buyer is the harder gap)
   On file: CT-1DB73E, Chief People Officer, economic buyer

7. Deal-885F45 (C-5E8EFB) — under-threaded (2 active < 3)
   Amount: not provided | Stage: not provided
   Active: 2 of 2 (economic buyer 2026-08-26; champion 2026-08-11)
   Personas present: economic buyer, champion | Missing: HR admin, IT security, finance
   Add: IT security (buyer+champion already covered; security review is the typical remaining gate among the missing three — stage data absent, so this is coverage-based)
   On file: CT-B3F25D, IT Security Lead, IT security

8. Deal-FCBE5B (C-737030) — single-threaded
   Amount: not provided | Stage: not provided
   Active: 1 of 1 (champion 2026-08-29)
   Personas present: champion | Missing: economic buyer, HR admin, IT security, finance
   Add: economic buyer
   On file: none on file

9. Deal-5408B0 (C-2AE3AA) — under-threaded (2 active < 3)
   Amount: not provided | Stage: not provided
   Active: 2 of 2 (champion 2026-09-01; HR admin 2026-08-18)
   Personas present: champion, HR admin | Missing: economic buyer, IT security, finance
   Add: economic buyer
   On file: CT-07FA76, Chief People Officer, economic buyer

10. Deal-C6D97A (C-5A8FC2) — under-threaded (all 3 active = one persona)
    Amount: not provided | Stage: not provided
    Active: 3 of 3 (champion ×3: 2026-08-31, 2026-08-20, 2026-08-05)
    Personas present: champion | Missing: economic buyer, HR admin, IT security, finance
    Add: economic buyer
    On file: none on file

11. Deal-F9A08A (C-0D15DF) — single-threaded
    Amount: not provided | Stage: not provided
    Active: 1 of 2 (champion 2026-09-03; CT-913581 economic buyer last engaged 2026-06-20 < 2026-07-26 → stale, excluded; not former, so re-engagement is an option)
    Personas present: champion | Missing: economic buyer, HR admin, IT security, finance
    Add: economic buyer (reactivate stale CT-913581 and/or work the new contact below)
    On file: CT-697541, Chief People Officer, economic buyer

NOT FLAGGED (for completeness): Deal-84DBA6 (3 active, 3 personas), Deal-4B0BEB (4 active, 4 personas), Deal-D348E1 (5 active, all 5 personas).

CAVEATS
- Amount and stage are absent from the provided data; every "most valuable persona" call is coverage-based, not stage-based.
- "Open" status is not in the data; all 14 deals were audited as given.
talk-track-from-transcriptsPASS 1.00
Plain-English question: From this rep's last 10 calls: how do they open, handle objections, and how often do they lock a next step?
Prompt — what the model saw
From this rep's last 10 calls: what they lead with in the first five minutes, how they handle the three most common objections, how often a concrete next step is agreed (rate), and every competitor a prospect raised. Quote at most one line per finding and only lines that appear in the transcripts. End with two coaching notes.
cost $0.0478 · 34s · in 3,479 / out 2,012 tok
✓ objection_count✓ next_step_rate✓ competitors✓ coaching✓ quote_used

✓ Correct Answer (ground truth)

Plain-English answer: 8 of 10 calls open with the same customer-win story; objection handling, next-step agreement rate, and every competitor raised, with quotes.
The answer in one line: n_transcripts = 10; objection_count = 3; rep = Alex Franklin (+3 more below)
n_transcripts10
objection_count3
allowed_new13 items — e.g. TT-001, TT-002, TT-003, TT-004
competitor_quotesAwardco=We're also in late talks with Awardco — their rewards catalog looks bigger than yours.; Kudos=How are you different from Kudos? Our CEO used them at her last company.
competitors_prospect_raisedAwardco, Kudos
next_stepagreed=7; total=10; rate=0.7
objectionsobjection=budget; count=4; quote=Honestly, budget is locked until next fiscal year — I can't add a new line item right now.; transcript_id=TT-001 | objection=timing; count=3; quote=This is good, but can we revisit it next quarter? Open enrollment starts in October and we're underwater.; transcript_id=TT-002 | objection=in-house; count=3; quote=We already do recognition with a spreadsheet and quarterly gift cards — why would we change?; transcript_id=TT-004
opener_themetheme=opens with a customer recognition win story (400-person retailer, 18% turnover reduction); count=8; total=10; quote=Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it.; transcript_id=TT-001
repAlex Franklin
rep_raised_competitors_excludedWorkhuman
graded checksobjection_count · next_step_rate · competitors · coaching · quote_used

Model's actual answer

show the model's full answer
**1) What Alex leads with in the first five minutes**

8 of 10 calls open with the identical retailer proof story at minute 0 (TT-001, TT-002, TT-003, TT-005, TT-006, TT-007, TT-008, TT-010). Quote (verbatim across all eight): "Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it."

The two exceptions: TT-004 opens with agenda-setting ("I put together a short agenda — security review first, then pricing.") and TT-009 opens on pricing ("You asked for straight pricing last time, so let's start there."). No discovery questions appear in any minute-0 opening in the data provided.

**2) Three most common objections and how they're handled**

Counts from prospect lines: budget locked = 4 (TT-001, TT-003, TT-006, TT-010, all min 6); "revisit next quarter" timing = 3 (TT-002, TT-005, TT-008, min 6); status quo spreadsheet/gift cards = 3 (TT-004, TT-007, TT-009, min 6). (Below the top three: committee gate = 2, TT-004/TT-010 min 11; no urgency = 1, TT-007 min 14.)

- Budget → reframes to self-funding via turnover savings. Quote: "Totally fair. Most teams fund this out of turnover savings — that retailer saved about $210k in avoided backfills, which is how their finance team signed off."
- Timing → offers a scoped pilot. Quote: "Makes sense. What if we scope a 90-day pilot with one department so you have internal data before next quarter's planning?"
- Status quo → contrasts automation and analytics. Quote: "Spreadsheets work until they scale — the difference is automation: milestones fire without HR lifting a finger, and you get analytics on who is being recognized."

All three rebuttals are word-for-word identical across calls. The two committee objections and the no-urgency deflection get no rebuttal: "Understood — I'll leave it with you." / "Understood, thanks for the candor." / "Fair enough."

**3) Concrete next-step agreement rate**

- Rep asked for a next step in 7 of 10 calls (TT-001, TT-002, TT-003, TT-005, TT-006, TT-008, TT-009 — "Should we lock the next step — a working session with your team this week?").
- All 7 asks were accepted: "Yes, Thursday at 2pm works — send the invite and I'll bring our HRIS manager."
- No ask in TT-004, TT-007, TT-010 — exactly the three calls that ended on committee/no-urgency deferrals.
- Rate: 7/10 = 70% of calls end with a concrete agreed next step (7/7 = 100% when the ask is made).

**4) Competitors raised by prospects**

- Awardco (TT-003, min 4): "We're also in late talks with Awardco — their rewards catalog looks bigger than yours."
- Kudos (TT-007, min 4): "How are you different from Kudos? Our CEO used them at her last company."

Workhuman appears only as a rep-initiated mention (TT-005, min 2), not raised by a prospect, so it's excluded from this list.

**Coaching notes**

1. The three lost-momentum calls (TT-004, TT-007, TT-010) share one pattern: when the prospect defers to a committee or says "no urgency," Alex accepts it and exits without any next step or offer to arm the internal champion. Even a fallback ask (e.g., sending materials for the committee with a dated follow-up) would lift the 70% rate; the data shows the 100% acceptance rate when an ask is actually made.
2. The opener and all three rebuttals are verbatim scripts (same retailer story 8/10, identical responses every time), and no discovery question appears in any opening in the transcripts. That leaves competitor moments (Awardco's catalog, Kudos's CEO relationship) answered with canned positioning rather than anything tailored to what the prospect just said.
reporting-analytics 5 tests
q3-weighted-forecastPASS 1.00
Plain-English question: Build the Q3 2026 weighted forecast from the open pipeline, and show the math deal by deal.
Prompt — what the model saw
Below are the open deals in the default sales pipeline with a close date between 2026-07-01 and 2026-10-15, pulled from the CRM warehouse on 2026-09-05. Stages: DS1 (discovery) through DS5 (contract). Build the Q3 2026 forecast (quarter = 2026-07-01 to 2026-09-30). Weighting rule: 100 percent of COMMIT plus 35 percent of BEST_CASE. PIPELINE deals count zero. Only deals with a close date inside the quarter count. Report: COMMIT total, BEST_CASE total, weighted forecast, count of deals in each category inside the quarter, the deals you excluded for being outside the quarter (count and total amount), and the top 5 BEST_CASE deals by amount inside the quarter. Show arithmetic. Then, in a final '## Data quality' section of no more than four sentences, name the data quality problems in this extract that would make the forecast unreliable if it were run unattended.
cost $0.2003 · 61s · in 29,696 / out 5,812 tok
✓ commit_total_44729✓ best_case_total_203565✓ weighted_115976.75✓ commit_count_7✓ best_case_count_24✓ excluded_count_32✓ excluded_amount_227575✓ top5_aliases✓ flags_blank_owner✓ flags_missing_why_buys

✓ Correct Answer (ground truth)

Plain-English answer: A stage-weighted Q3 forecast number computed from the open pipeline, with per-stage math that reconciles to the deal data.
The answer in one line: excluded_amount = 227575; best_case_total = 203565; weighted_forecast = 115,976.75; commit_total = 44729; excluded_out_of_quarter = 32; best_case_count = 24; flags = blank owner column + empty why_buys fields must be surfaced (+2 more below)
excluded_amount227,575
best_case_count24
best_case_total203,565
commit_count7
commit_total44,729
excluded_out_of_quarter32
weighted_forecast115,976.75
flagsblank owner column + empty why_buys fields must be surfaced
top5_best_caseDeal-2D7423, Deal-25F752, Deal-E53952, Deal-5EED42, Deal-FA32A0
graded checkscommit_total_44729 · best_case_total_203565 · weighted_115976.75 · commit_count_7 · best_case_count_24 · excluded_count_32 · excluded_amount_227575 · top5_aliases · flags_blank_owner · flags_missing_why_buys

Model's actual answer

show the model's full answer
Q3 2026 FORECAST (quarter = 2026-07-01 to 2026-09-30; extract of 86 deals pulled 2026-09-05)

DEALS INSIDE THE QUARTER: 54 of 86

COMMIT (in-quarter): 7 deals, total $44,729
  Deal-547B2B (DS5)  11,200
  Deal-B7EBD1 (DS5)   9,000
  Deal-403845 (DS5)   9,000
  Deal-A2B47C (DS5)   6,360
  Deal-2465CE (DS5)   5,400
  Deal-A5E80A (DS1)   2,520
  Deal-499BF6 (DS2)   1,249
  Check: 11,200+9,000+9,000+6,360+5,400+2,520+1,249 = 44,729

BEST_CASE (in-quarter): 24 deals, total $203,565
  (largest: Deal-2D7423 38,935; Deal-25F752 24,000; Deal-E53952 19,656; ... smallest: Deal-87412C 528)

PIPELINE (in-quarter): 23 deals, total $201,637.40 — weighted at $0 per rule (Deal-42326B carries the non-integer $2,480.40)

WEIGHTED FORECAST
  = 100% × COMMIT + 35% × BEST_CASE + 0% × PIPELINE
  = 44,729 + (0.35 × 203,565) + 0
  = 44,729 + 71,247.75
  = $115,976.75

EXCLUDED — CLOSE DATE OUTSIDE QUARTER: 32 deals, total $227,575
  All 32 have close dates in October 2026 (2026-10-01 through 2026-10-15); none fall before 2026-07-01. Breakdown: 22 PIPELINE, 9 BEST_CASE, 1 COMMIT (Deal-D348E1, $13,770, 2026-10-15).

TOP 5 BEST_CASE DEALS BY AMOUNT (in-quarter)
  1. Deal-2D7423 (DS3)  $38,935  close 2026-09-30
  2. Deal-25F752 (DS4)  $24,000  close 2026-09-25
  3. Deal-E53952 (DS4)  $19,656  close 2026-09-30
  4. Deal-5EED42 (DS3)  $16,250  close 2026-09-30
  5. Deal-FA32A0 (DS3)  $11,116  close 2026-09-25

## Data quality
85 of 86 deals have a blank owner (only Deal-C9C286 names Bryce Harmon), so nothing can be attributed, rolled up by rep, or chased for a forecast submission. Forecast categories contradict stage: two COMMIT deals sit at DS1/DS2 (Deal-A5E80A, Deal-499BF6) where commit is implausible, and one DS5 contract-stage deal is categorized BEST_CASE (Deal-C61CF7), so the weighting is being applied to labels that don't match pipeline reality. 71 of 86 deals have why_buys_chars = 0, meaning most of the forecasted value has no documented buying rationale behind it. The extract also stops at 2026-10-15 while Q4 deals beyond that window may exist, and there are no company/account names — only aliases — so duplicates or split deals can't be detected.
aha-moment-2x20.67
Plain-English question: Do companies that hit 5+ givers and one redemption in month one actually retain better at 24 months?
Prompt — what the model saw
Test the activation hypothesis on the company cohort below: companies with 5+ unique givers (m1_users >= 5) AND 1+ successful redemption (m1_redemptions >= 1) in their first calendar month retain better at 24 months than those with one signal or neither. Every company in this extract is already 25+ months old, so current_status = 'active' means retained at 24 months. Report the full 2x2 (cohort size and 24-month retention rate for: both signals, givers-only, redemption-only, neither), state how many companies were excluded from the denominator and why (if none, say so), name the single signal with the largest retention lift, and state what this does and does not prove.
cost $0.2746 · 131s · in 23,455 / out 12,976 tok
✓ both_66✓ givers_47✓ redemp_31✓ neither_40✗ excluded_0✗ largest_lift_givers

✓ Correct Answer (ground truth)

Plain-English answer: The four cohort cells are graded against computed 24-month retention; the numbers come from the scorer's cohort table, not opinion.
The answer in one line: both_5plus_givers_and_redemption = 66.0% (31 of 47)
both_5plus_givers_and_redemption66.0% (31 of 47)
exclusionsnone — claim bounded as correlation, not causation
givers_only46.9% (23 of 49)
largest_single_signal_liftunique givers
neither40.0% (38 of 95)
redemption_only31.0% (9 of 29)
graded checksboth_66 · givers_47 · redemp_31 · neither_40 · excluded_0 · largest_lift_givers

Model's actual answer

show the model's full answer
ACTIVATION HYPOTHESIS TEST — plg_company_cohort_2023.csv

Definitions applied exactly as specified:
- Signal G ("givers"): m1_users >= 5
- Signal R ("redemption"): m1_redemptions >= 1
- Retained at 24 months: current_status = 'active'. 'cancelled' and 'non_renewing' both count as NOT retained.

DENOMINATOR / EXCLUSIONS
All 220 rows are included. Zero companies excluded. Signup months run 2023-01 through 2023-07, so every company is 25+ months old as stated; no missing/invalid values in the four fields used (m1_users, m1_redemptions, current_status). Status mix: 101 active, 116 cancelled, 3 non_renewing.

FULL 2x2

```
Cell               Cohort (n)   Active   24-mo retention   Arithmetic
Both signals (G+R)     47         31        66.0%          31/47 = 0.6596
Givers-only (G,¬R)     49         23        46.9%          23/49 = 0.4694
Redemption-only (¬G,R) 29          9        31.0%           9/29 = 0.3103
Neither (¬G,¬R)        95         38        40.0%          38/95 = 0.4000
TOTAL                 220        101        45.9%         101/220
```

Lifts vs "neither" baseline (40.0%):
- Both: 66.0% − 40.0% = +26.0 pp
- Givers-only: 46.9% − 40.0% = +6.9 pp
- Redemption-only: 31.0% − 40.0% = −9.0 pp (BELOW baseline)

Marginal (single-signal) lifts across the whole cohort:
- m1_users >= 5: 54/96 = 56.2% vs 47/124 = 37.9% → +18.4 pp
- m1_redemptions >= 1: 40/76 = 52.6% vs 61/144 = 42.4% → +10.2 pp

LARGEST SINGLE-SIGNAL LIFT
Givers (m1_users >= 5), at +18.4 pp marginal vs redemption's +10.2 pp. It is also the only single signal that beats the neither-baseline on its own (+6.9 pp); redemption-only actually retains worse than neither (−9.0 pp).

WHAT THIS PROVES
- In this 220-company 2023 cohort, hitting both signals in month 1 is associated with materially higher 24-month retention (66.0%) than either signal alone or neither.
- The association is strongest and most consistent for early giver breadth; early redemption alone is not a positive retention signal here (31.0%, and its marginal lift is driven entirely by companies that also have the giver signal).

WHAT THIS DOES NOT PROVE
- Causation. This is observational; activated companies may differ systematically (size, tier, industry) from non-activated ones. Note the confound visible in the data itself: all 3 non_renewing and 116 cancelled rows include many tier_two/cancelled-with-$0-ARR companies, and tier correlates with outcome — activation may be a marker, not a driver.
- The "both > either" pattern is consistent with a threshold/interaction effect, but no significance test was requested or run; cells of n=29–49 are small enough that a few companies shift rates by several points.
- Nothing about cohorts outside 2023-01→2023-07 signups, about m3 behavior, or about ARR/expansion — retention here is binary status only.
arr-rollforward-reconciliation0.80
Plain-English question: Does CRM ARR match billing ARR — and if not, exactly where is the variance?
Prompt — what the model saw
Reconcile ARR between CRM company records and billing subscriptions as of 2026-09-05. Billing ARR is MRR x 12. Report both totals, the variance, and a decomposition into buckets (status mismatch, rounding, missing records, other) that sums exactly to the variance. List each mismatched account with a suggested owner. Business rule: any subscription with a term other than 12 months must have cf_agreement_end_date populated; list violations.
cost $0.1084 · 70s · in 5,655 / out 5,442 tok
✓ billing_total✓ crm_total✗ variance✓ buckets_sum✓ violations_listed

✓ Correct Answer (ground truth)

Plain-English answer: Billing $604,739.28 vs CRM $603,581.76: a $1,157.52 variance decomposed into status-mismatch, rounding, and missing-record buckets.
The answer in one line: n_companies = 40; n_mismatched_accounts = 7; as_of = 2026-09-05 (+4 more below)
n_companies40
n_mismatched_accounts7
as_of2026-09-05
billing_arr_total604739.28
bucket_sum_equals_variance1
buckets{"status_mismatch": {"amount": "-13158.48", "accounts": [{"company_alias": "C-0C8323BF", "billing_arr": "0.00", "hubspot_arr": "4905.24", "diff": "-4905.24", "suggested_owner": "Dana Mercer"}, {"compa
crm_arr_total603581.76
ruleterm_months != 12 requires cf_agreement_end_date
variance1157.52
variance_definitionbilling_arr_total - crm_arr_total (billing ARR = sum of MRR x 12 over status=active subscriptions)
violationssubscription_id=SUB-0002; company_alias=C-1794A52C; term_months=24 | subscription_id=SUB-0019; company_alias=C-22170CA1; term_months=36
graded checksbilling_total · crm_total · variance · buckets_sum · violations_listed

Model's actual answer

show the model's full answer
ARR RECONCILIATION — as of 2026-09-05

TOTALS
- Billing ARR (all 40 subscriptions, MRR x 12): $617,897.76
  (includes 2 cancelled subs totaling $13,158.48; active-only billing = $604,739.28)
- CRM total (40 company records, hubspot_arr): $603,581.76
- Variance (CRM - Billing): $603,581.76 - $617,897.76 = -$14,316.00

DECOMPOSITION (sums exactly to -$14,316.00)

1. Status mismatch: $0.00
   The two cancelled subs — SUB-000E (C-0C8323BF, $4,905.24) and SUB-000F (C-0DC4FB8C, $8,253.24) — are included in the billing total, and CRM carries exactly matching ARR ($4,905.24 and $8,253.24). Net contribution to variance: zero. Flagged separately because CRM appears to count cancelled accounts as ARR.

2. Rounding: +$36.00
   - C-0D66DF9E: CRM $23,200.00 - billing $23,184.00 (1,932.00 x 12) = +$16.00
   - C-14D70CE0: CRM $18,200.00 - billing $18,180.00 (1,515.00 x 12) = +$20.00

3. Missing records: -$11,952.00
   - C-21629AA4 (SUB-0004, active, $28,449.24 = 2,370.77 x 12): no CRM record -> -$28,449.24
   - C-0D5BBE3A (CRM $16,497.24): no billing subscription -> +$16,497.24
   Net: -$28,449.24 + $16,497.24 = -$11,952.00

4. Other (unexplained data mismatch): -$2,400.00
   - C-0F7269D7 (SUB-0006, active): billing $26,796.00 (2,233.00 x 12) vs CRM $24,396.00. The $2,400.00 gap is not explainable from the provided data (could be an unbooked price increase, but no evidence given).

Check: $0.00 + $36.00 - $11,952.00 - $2,400.00 = -$14,316.00 ✓

MISMATCHED ACCOUNTS & SUGGESTED OWNER
Note: no owner/rep field exists in either file, so owners cannot be identified from the data. Suggested owner below is a routing recommendation, not a data-derived fact.

- C-21629AA4 — billing $28,449.24, CRM record missing. Suggested owner: RevOps (create/sync company record).
- C-0D5BBE3A — CRM $16,497.24, no active or cancelled subscription. Suggested owner: RevOps + Billing ops (verify churn not recorded in billing, or stale CRM ARR).
- C-0F7269D7 — billing $26,796.00 vs CRM $24,396.00 (-$2,400.00). Suggested owner: RevOps (CRM ARR field update).
- C-0C8323BF, C-0DC4FB8C — cancelled in billing, still carrying ARR in CRM. Suggested owner: RevOps (confirm whether CRM ARR should be zeroed).
- C-0D66DF9E (+$16.00), C-14D70CE0 (+$20.00) — immaterial rounding; suggest leave as-is or RevOps tidy-up.

BUSINESS RULE VIOLATIONS (term != 12 months requires cf_agreement_end_date)
- SUB-0002 — C-1794A52C, term 24 months, active, cf_agreement_end_date EMPTY. Violation.
- SUB-0019 — C-22170CA1, term 36 months, active, cf_agreement_end_date EMPTY. Violation.

Compliant non-12-month subs: SUB-000C (C-0DB48281, 24mo, end 2027-11-30) and SUB-001A (C-0FC4DBB8, 36mo, end 2027-11-30).
kvm-month-over-monthPASS 1.00
Plain-English question: How did the four core KVMs move in August versus July?
Prompt — what the model saw
Report the four core KVMs for 2026-08 against 2026-07: giving rate, redemptions per user, 1:1 meetings engagement, pulse check engagement. For each: value, prior value, absolute and relative change, direction. Name the KVM with the largest relative move and the segment (plan_tier or size_band) driving it if the data supports that; otherwise say the data does not. Under 200 words plus one table.
cost $0.1061 · 54s · in 7,918 / out 4,452 tok
✓ four_kvms✓ biggest_mover_named✓ direction_correct✓ segment_named✓ word_cap

✓ Correct Answer (ground truth)

Plain-English answer: Giving rate 60.27% (up 0.07%), redemptions per user 1.7302 (up 0.01%), with 1:1 meetings and pulse engagement each up fractionally — direction reported for all four.
The answer in one line: n_company_rows_per_month = 30; aggregation = simple mean of company-level values per month
n_company_rows_per_month30
aggregationsimple mean of company-level values per month
biggest_moverpulse_engagement
driving_segmentsize_band=enterprise; 2026-07=0.55; 2026-08=0.2743; rel_change_pct=-50.13
kvms{"giving_rate": {"2026-07": 0.6023, "2026-08": 0.6027, "abs_change": 0.0004, "rel_change_pct": 0.07, "direction": "up"}, "redemptions_per_user": {"2026-07": 1.73, "2026-08": 1.7302, "abs_change": 0.00
segment_detail{"smb": {"2026-07": 0.6588, "2026-08": 0.6573, "rel_change_pct": -0.23}, "mid_market": {"2026-07": 0.593, "2026-08": 0.5942, "rel_change_pct": 0.2}, "enterprise": {"2026-07": 0.55, "2026-08": 0.2743,
graded checksfour_kvms · biggest_mover_named · direction_correct · segment_named · word_cap

Model's actual answer

show the model's full answer
KVM MoM report, 2026-08 vs 2026-07. Method: unweighted mean across the 30 companies (identical cohort both months; all plan_tier=tier_three, so tier gives no segmentation). 30 rows/month.

| KVM | 2026-08 | 2026-07 | Abs Δ | Rel Δ | Direction |
|---|---|---|---|---|---|
| Giving rate | 0.6027 | 0.6023 | +0.0004 | +0.07% | Up (flat) |
| Redemptions per user | 1.7302 | 1.7300 | +0.0002 | +0.01% | Up (flat) |
| 1:1 meetings engagement | 0.4472 | 0.4469 | +0.0003 | +0.06% | Up (flat) |
| Pulse check engagement | 0.5086 | 0.6006 | −0.0920 | −15.31% | Down |

Arithmetic example (pulse): Jul mean = Σ/30 = 0.600587; Aug mean = 0.508610; Δ = −0.091977; rel = −0.091977/0.600587 = −15.31%.

Largest relative move: pulse check engagement, −15.31%. The data supports the driving segment: size_band=enterprise. Enterprise pulse fell from 0.5500 to 0.2743 (−50.13%), while mid_market (+0.21%) and smb (−0.22%) were flat. All 10 enterprise companies (C-0B2895EF, C-0B2213A9, C-0D6CC8E3, C-0D0B047C, C-0D3278C7, C-0FCCD2DF, C-0F6C0F34, C-8C2E8F00, C-0B827671, C-0BA71F12) roughly halved — a uniform, cohort-wide pattern suggesting a tracking/instrumentation change rather than genuine behavior decline; the file contains no field to confirm cause. plan_tier cannot drive anything (single tier).
rewards-ytd-sectionPASS 1.00
Plain-English question: Produce the redemption section of the weekly rewards report through August.
Prompt — what the model saw
Produce the redemption section of the weekly rewards report through the last completed month (2026-08): redemption count, spend, unique redeemers, redemptions per redeemer, provider mix as percent of spend, top 5 countries by redemptions. Say which month is the last completed month. Provider shares must sum to 100.
cost $0.5836 · 241s · in 52,681 / out 23,564 tok
✓ count✓ spend✓ unique_redeemers✓ month_named✓ provider_mix✓ top_country

✓ Correct Answer (ground truth)

Plain-English answer: 378 redemptions, $27,846 spend, 235 unique redeemers at 1.61 each, custom rewards at 39% of spend, US leading all countries.
The answer in one line: redemption_count = 378; unique_redeemers = 235; redemptions_per_redeemer = 1.61; total_spend_usd = 27846.00
redemption_count378
redemptions_per_redeemer1.61
unique_redeemers235
provider_mix_pct_of_spendTangoCard=11.6; Tremendous=30.54; custom=39.05; Snappy=18.81
top5_countries_by_countcountry=US; count=244 | country=CA; count=24 | country=AU; count=21 | country=GB; count=17 | country=NL; count=17
total_spend_usd27846.00
graded checkscount · spend · unique_redeemers · month_named · provider_mix · top_country

Model's actual answer

show the model's full answer
REDEMPTION SECTION — YTD Jan 1 – Aug 31, 2026

Last completed month: 2026-08 (August 2026). All 378 rows in redemptions_ytd.csv fall within Jan–Aug 2026; none excluded.

Headline metrics
- Redemptions: 378
- Spend: $27,846.00 (sum of amount_usd across all 378 rows)
- Unique redeemers: 235 (distinct user_key, deduped — not a sum of monthly figures)
- Redemptions per redeemer: 378 ÷ 235 = 1.61

Provider mix (% of spend) — shares forced to sum to 100.0 via largest-remainder adjustment (+0.1pp to custom, the largest fractional remainder):
- custom: $10,873.00 → 39.0469% → 39.1% (37 redemptions)
- Tremendous: $8,505.00 → 30.5430% → 30.5% (192 redemptions)
- Snappy: $5,238.00 → 18.8106% → 18.8% (59 redemptions)
- TangoCard: $3,230.00 → 11.6002% → 11.6% (90 redemptions)
- Check: 39.1 + 30.5 + 18.8 + 11.6 = 100.0 ✓; spend check: 10,873 + 8,505 + 5,238 + 3,230 = $27,846 ✓; volume check: 37 + 192 + 59 + 90 = 378 ✓

Top 5 countries by redemptions
1. US — 244
2. CA — 24
3. AU — 21
4. GB — 17 (tied)
4. NL — 17 (tied)

Note: GB and NL tie at 17 redemptions for the 4th/5th slots. (Next: SG at 12.)

Provider names are cited exactly as they appear in the data ("custom" is lowercase as given). All figures computed from the provided CSV only; no external data used.
customer-success 5 tests
churn-save-eligibilityPASS 1.00
Plain-English question: Which at-risk accounts actually qualify for a churn-save offer, and which play fits each?
Prompt — what the model saw
Which at-risk accounts qualify for a churn-save offer under the documented eligibility rules, what amount is at stake per account and in total, and which play fits each (usage revival, executive touch, commercial concession)? Cite the signal that justifies each play. List accounts that look at risk but do not qualify and why.
cost $0.0537 · 43s · in 2,024 / out 2,987 tok
✓ eligible_set✓ total_at_stake✓ plays_cited✓ noneligible_named✓ rules_applied

✓ Correct Answer (ground truth)

Plain-English answer: Only accounts meeting all three rules qualify — e.g. C-0F6C0F34 at $49,707 (executive touch) and C-0B827671 at $25,365 (usage revival) — with amount and play per account.
The answer in one line: n_accounts = 30; snapshot_date = 2026-09-05
n_accounts30
eligibleaccount_alias=C-0F6C0F34; amount_at_stake=49707.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0B827671; amount_at_stake=25365.00; play=usage revival; justifying_signal=usage_trend_3m=declining over the last 3 months | account_alias=C-0B360C78; amount_at_stake=35748.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 75% - risk is commercial, not adoption | account_alias=C-0B0F1BAB; amount_at_stake=5494.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0CA21961; amount_at_stake=16829.00; play=usage revival; justifying_signal=seat utilization 26% is below 50% | account_alias=C-0E9C27D1; amount_at_stake=41235.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 85% - risk is commercial, not adoption … (+2 more)
non_eligible_at_riskaccount_alias=C-0BC71BDD; health_score=55 | account_alias=C-0BA71F12; health_score=52 | account_alias=C-0F6694C3; health_score=43 | account_alias=C-0BE96399; health_score=54 | account_alias=C-0F876796; health_score=47 | account_alias=C-0FCCD2DF; health_score=43 … (+1 more)
ruleshealth_score < 60, churn_save_eligible_amount > 0, renewal within 120 days of snapshot
snapshot_date2026-09-05
total_amount_at_stake224601.00
graded checkseligible_set · total_at_stake · plays_cited · noneligible_named · rules_applied

Model's actual answer

show the model's full answer
QUALIFIED ACCOUNTS (all 3 rules must pass)

Eligibility arithmetic:
- R1: health_score < 60
- R2: churn_save_eligible_amount > 0
- R3: renewal_date within 120 days of 2026-09-05 → cutoff = 2027-01-03 (Sep 25 + Oct 31 + Nov 30 + Dec 31 = 117 days to Dec 31; +3 = Jan 3)

Play assignment logic (from available signals): usage_trend declining or seat utilization <50% → usage revival; champion_active = false → executive touch; engagement signals healthy but low health + near renewal with no usage/champion lever → commercial concession.

8 accounts qualify:

```
Account     Health  At-stake     Renewal     Play                  Justifying signal
C-0F6C0F34  51      $49,707      2026-10-03  Executive touch       champion_active=false (usage growing, 308/395=78% util)
C-0B827671  56      $25,365      2026-11-14  Usage revival         usage_trend_3m=declining (113/202=56% util)
C-0B360C78  57      $35,748      2026-10-28  Commercial concession usage growing + champion active + 75% util → no engagement lever; health 57, renewal in 53 days
C-0B0F1BAB  38      $5,494       2026-09-23  Executive touch       champion_active=false; renewal in 18 days (flat usage, 66% util)
C-0CA21961  58      $16,829      2026-12-28  Usage revival         84/325 seats used = 26% utilization (flat trend, champion active)
C-0E9C27D1  39      $41,235      2026-09-24  Commercial concession 85% util + champion active + flat usage → engagement is fine; health 39, renewal in 19 days
C-0CEF69FD  53      $32,621      2026-11-21  Executive touch       champion_active=false (usage growing, 97/136=71% util)
C-0D3278C7  54      $17,602      2026-11-12  Usage revival         usage_trend_3m=declining (126/380=33% util)
```

Total at stake: 49,707 + 25,365 + 35,748 + 5,494 + 16,829 + 41,235 + 32,621 + 17,602 = $224,601

AT-RISK BUT NOT QUALIFIED (health < 60, fails R2 or R3)

```
Account     Health  Fails  Why
C-0BA71F12  52      R3     renewal 2027-04-11 = 218 days out (>120)
C-0F876796  47      R3     renewal 2027-02-06 = 154 days out (>120)
C-0BC71BDD  55      R2     churn_save_eligible_amount = 0
C-0F6694C3  43      R2     churn_save_eligible_amount = 0
C-0BE96399  54      R2     churn_save_eligible_amount = 0
C-0FCCD2DF  43      R2     churn_save_eligible_amount = 0
C-10A56B0F  54      R2     churn_save_eligible_amount = 0
```

Note on missing data: the file provides no explicit play-assignment rules, so plays above are inferred from the documented signal fields (usage_trend_3m, seats_used/seats, champion_active) using the stated logic. All other accounts in the file (health ≥ 60) are not at risk under R1.
expansion-kitPASS 1.00
Plain-English question: Build the expansion kit for the customer who just replied to our outreach.
Prompt — what the model saw
A customer replied to outreach (reply included). Build the expansion kit: seat coverage (licensed over headcount), two lines on usage health, headroom in seats and ARR at the current per-seat rate, who replied and whether they can buy, the right buyer from our contacts, and a reply email under 150 words that moves toward the right person without being pushy and cites one real usage fact.
cost $0.0755 · 34s · in 9,574 / out 1,479 tok
✓ coverage_pct✓ headroom_arr✓ right_buyer✓ replier_not_buyer✓ usage_fact✓ email_under_150

✓ Correct Answer (ground truth)

Plain-English answer: 37.5% seat coverage, 250 headroom seats worth $15,000 ARR at $60 a seat, usage up 43.2% in six months — and the replier cannot buy; Dana R. (VP People) is the right buyer.
The answer in one line: headroom_seats = 250; coverage_pct = 37.5; company_alias = C-0DDFC9A7 (+3 more below)
headroom_seats250
coverage_pct37.5
company_aliasC-0DDFC9A7
headroom_arr15000.00
per_seat_annual_rate60.00
replierMaria S.
replier_can_buy0
right_buyername=Dana R.; title=VP People
usage_factsactive_users_2026_08=126; active_users_2026_03=88; growth_pct_6m=43.2; licensed_seat_utilization_pct=84
usage_health_linesMonthly active users grew from 88 (2026-03) to 126 (2026-08), +43.2% over 6 months., Last month 126 of 150 licensed seats were active (84.0% utilization), approaching the seat cap.
graded checkscoverage_pct · headroom_arr · right_buyer · replier_not_buyer · usage_fact · email_under_150

Model's actual answer

show the model's full answer
EXPANSION KIT — C-0DDFC9A7

1) SEAT COVERAGE
Licensed / headcount = 150 / 400 = 37.5% of headcount licensed.

2) USAGE HEALTH (two lines)
Monthly active users rose every month Mar–Aug 2026: 88 → 95 → 102 → 110 → 118 → 126 (+38 users, avg +7.6/month, no down months).
Aug actives (126) = 84% of the 150 licensed seats (126/150 = 0.84) — adoption is high and still climbing toward the license cap.

3) HEADROOM
Seats: 400 − 150 = 250 unlicensed seats.
Per-seat rate: $9,000 ARR / 150 seats = $60/seat/year.
ARR headroom: 250 × $60 = $15,000 (i.e., 2.67x current ARR; total at full coverage = $24,000).

4) WHO REPLIED / CAN THEY BUY
Maria S., People Operations Coordinator (last engaged 2026-09-02). No — she states purchasing and seat expansion sit with Dana R., and offered an introduction.

5) RIGHT BUYER (from contacts)
Dana R., VP People — confirmed budget/seat-expansion owner per Maria's reply; last engaged 2026-05-18. Note: she's been asking about usage numbers (per Maria), so usage data is the right opener. Sam K. (Office Manager, last engaged 2025-11-03) is not indicated as a buyer.

6) REPLY EMAIL (118 words)

Subject: RE: Growing your team's recognition program

Hi Maria,

Thanks so much for the quick reply — and glad to hear the team loves Bonusly and the feed is busy every week. The numbers back that up: monthly active users have climbed from 88 in March to 126 in August, a 43% increase in six months.

Since Dana owns budget and seat expansion, an introduction would be genuinely helpful — especially with only 150 seats licensed against a 400-person team. If it's easier, I'm happy to send over a short usage summary she can review before we talk.

Really appreciate you pointing me in the right direction.

Best,
Cole

Missing data note: no pricing/discount terms, renewal date, or Dana's direct contact info were provided; the per-seat rate above is derived from current ARR ÷ licensed seats only.
onboarding-statusPASS 1.00
Plain-English question: What is done and what is missing in this customer's onboarding before the mid-onboarding call?
Prompt — what the model saw
Prep the CSM for the mid-onboarding call: what is complete (integrations connected, allowance set, admins added, first recognition given, first redemption), what is not, early engagement signals, and the three things to cover on the call. Mark nothing complete without a data field that shows it.
cost $0.0268 · 22s · in 829 / out 1,147 tok
✓ complete_items✓ incomplete_items✓ three_items✓ grounded

✓ Correct Answer (ground truth)

Plain-English answer: 25 days live: Slack connected, allowance set, admins added — but the HRIS integration is not connected, with early engagement signals and three call priorities.
The answer in one line: days_live = 25; company_alias = C-0D284E42
days_live25
call_itemsConnect the HRIS integration (integration_hris is empty) to automate user provisioning and offboarding., Get to first redemption (first_redemption_at is empty): walk admins through the rewards catalog and confirm allowance funding., Build the giving habit: daily active givers averaged 4.3 in week 1 vs 13.0 last week - agree on manager nudges and a recognition cadence.
checklist{"slack_integration_connected": {"complete": true, "field": "integration_slack", "value": "2026-08-12"}, "hris_integration_connected": {"complete": false, "field": "integration_hris", "value": ""}, "a
company_aliasC-0D284E42
engagement_signalsavg_active_givers_week1=4.3; avg_active_givers_last_week=13; latest_active_givers=15; days_with_activity=25
graded checkscomplete_items · incomplete_items · three_items · grounded

Model's actual answer

show the model's full answer
CSM PREP — MID-ONBOARDING CALL
Account: C-0D284E42 | Signup: 2026-08-11 | Data window: 2026-08-11 → 2026-09-04 (24 days)

COMPLETE (data field present)
- Slack integration connected: 2026-08-12 (day 1 post-signup)
- Allowance set: 2026-08-13 (day 2)
- Admins added: 2
- First recognition given: 2026-08-15 14:22 (day 4)

NOT COMPLETE (field empty — nothing marked without data)
- HRIS integration: integration_hris is blank → not connected
- First redemption: first_redemption_at is blank → has not happened as of 2026-09-04

EARLY ENGAGEMENT SIGNALS (active_givers)
- Week 1 (08-11→08-17): 3+3+4+4+5+4+7 = 30 → avg 30/7 ≈ 4.3 givers/day
- Last 7 days (08-29→09-04): 11+13+11+13+13+15+15 = 91 → avg 91/7 = 13.0 givers/day
- Growth: 13.0 / 4.3 ≈ 3.0x week-over-week-average; day-1 value 3 → day-24 value 15 = 5x
- Trend is monotonically upward with no drop-off days; strongest signal in the dataset
- Caveat: no headcount/employee-count field exists, so penetration rate cannot be computed — absolute counts only

THREE THINGS TO COVER ON THE CALL
1. HRIS integration — the only setup step still missing. Get it connected; without it, eligibility/allowance sync likely stays manual.
2. First redemption — recognitions are flowing (first on day 4) but no redemption in 24 days. Walk through the catalog and reward options so points convert; non-redeemers churn risk.
3. Convert giver momentum into breadth — givers tripled their weekly average (4.3 → 13.0/day). Discuss expanding beyond the 2 current admins and formalizing recognition moments (channels, cadence) to sustain the curve.

MISSING DATA (explicit)
- HRIS connection date, first redemption date, company size/headcount, recognition volume, allowance amount — none provided; no assumptions made.
renewal-risk-conflicting-datesPASS 1.00
Plain-English question: Two systems disagree on renewal dates — which do you trust per account, and who is actually renewing inside 90 days?
Prompt — what the model saw
Produce a 90-day renewal risk brief. Two systems hold renewal dates and disagree on some accounts; decide which to trust per account and say why (multi-year contracts are known to be wrong in ChurnZero). For every renewal: company, CSM, ARR, date used, seat utilization, 3-month usage trend, risk rating with one sentence of evidence. Flag every disagreement. Close with total ARR renewing and ARR at risk.
cost $0.1906 · 115s · in 13,387 / out 9,547 tok
✓ total_renewing✓ arr_at_risk✓ disagreements_flagged✓ trust_rule

✓ Correct Answer (ground truth)

Plain-English answer: Chargebee wins per account (multi-year contracts are known wrong in ChurnZero), with each trusted date, the reason, and a 90-day-window flag.
The answer in one line: n_accounts = 20; n_disagreements = 5; snapshot_date = 2026-09-05 (+2 more below)
n_accounts20
n_disagreements5
accounts20 items — e.g. account_alias=C-0B144C78; csm=Cole Ingram; arr=30899.00; trusted_renewal_date=2026-11-02; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.4; usage_3m_ratio=1.03; risk=low; evidence=3-month usage ratio 1.03 (last3 avg 103 vs prior3 100), seat utilization 75% | account_alias=C-0B20DB64; csm=Dana Mercer; arr=21770.00; trusted_renewal_date=2026-10-07; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=56.6; usage_3m_ratio=1; risk=medium; evidence=3-month usage ratio 1.00 (last3 avg 295 vs prior3 295), seat utilization 57% | account_alias=C-0B344485; csm=Elena Sinclair; arr=64384.00; trusted_renewal_date=2026-11-16; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=78; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 241 vs prior3 231), seat utilization 78% | account_alias=C-0B369871; csm=Dana Mercer; arr=85128.00; trusted_renewal_date=2026-10-29; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.1; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 330 vs prior3 316), seat utilization 75%
arr_at_risk_90d359409.00
disagreementsaccount_alias=C-0B7D2C30; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-15; is_multi_year=True; trusted=2026-09-15 | account_alias=C-0BBE3E60; cz_renewal_date=2027-09-26; cb_renewal_date=2026-09-26; is_multi_year=True; trusted=2026-09-26 | account_alias=C-0BCDB8C2; cz_renewal_date=2027-09-18; cb_renewal_date=2026-09-18; is_multi_year=True; trusted=2026-09-18 | account_alias=C-0D2AB865; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-22; is_multi_year=True; trusted=2026-09-22 | account_alias=C-0F5D2323; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-29; is_multi_year=True; trusted=2026-09-29
snapshot_date2026-09-05
total_arr_renewing_90d1048715.00
trust_rulemulti-year contracts: Chargebee is authoritative (ChurnZero known wrong); otherwise systems agree or Chargebee wins
window2026-09-05 to 2026-12-04
graded checkstotal_renewing · arr_at_risk · disagreements_flagged · trust_rule

Model's actual answer

show the model's full answer
90-DAY RENEWAL RISK BRIEF
Window: 2026-09-24 through 2026-12-23. All 20 accounts in both files fall inside the window (three renewal dates have already passed — flagged below). No data was missing for any account.

SOURCE-OF-TRUTH DECISION
Rule applied: where ChurnZero (CZ) and Chargebee (CB) disagree AND CB marks the contract multi-year (is_multi_year=true), trust Chargebee — per the known issue that multi-year contracts are wrong in ChurnZero. All 5 disagreements are exactly the 5 multi-year accounts, so the rule resolves cleanly. For the 15 single-year accounts, CZ and CB dates match to the day — no conflict.

DISAGREEMENTS FLAGGED (all 5)
- C-0B7D2C30: CZ 2026-09-10 vs CB 2026-09-15 (36mo) → using 2026-09-15 (already passed; -9 days from today)
- C-0BCDB8C2: CZ 2027-09-18 vs CB 2026-09-18 (36mo) → using 2026-09-18 (passed; CZ shows the annual anniversary, not the true term date)
- C-0D2AB865: CZ 2026-09-10 vs CB 2026-09-22 (24mo) → using 2026-09-22 (passed)
- C-0BBE3E60: CZ 2027-09-26 vs CB 2026-09-26 (24mo) → using 2026-09-26 (2 days out)
- C-0F5D2323: CZ 2026-09-10 vs CB 2026-09-29 (24mo) → using 2026-09-29 (5 days out)

METHOD (arithmetic shown per account)
- Utilization = seats_used / seats.
- 3-month usage trend = avg(Jun+Jul+Aug 2026) vs avg(Mar+Apr+May 2026), % change.
- Risk rubric: HIGH = trend ≤ -10% OR utilization < 30%. MEDIUM = trend between -10% and 0% OR utilization 30-60%. LOW = trend ≥ 0% AND utilization ≥ 60%.

RENEWALS (sorted by date used)

HIGH RISK — $359,409 total
1. C-0B7D2C30 | Dana Mercer | $65,901 | 2026-09-15 (CB, PASSED) | util 274/476 = 57.6% | trend: (97+94+84)/3=91.7 vs (119+110+107)/3=112.0 = -18.2% | HIGH: 12 straight months of user decline (155→84).
2. C-0BCDB8C2 | Cole Ingram | $54,427 | 2026-09-18 (CB, PASSED) | util 232/424 = 54.7% | trend: (127+118+110)/3=118.3 vs (152+143+136)/3=143.7 = -17.6% | HIGH: every month down 12 months running (200→110).
3. C-0D2AB865 | Elena Sinclair | $38,022 | 2026-09-22 (CB, PASSED) | util 250/407 = 61.4% | trend: (125+117+109)/3=117.0 vs (152+144+137)/3=144.3 = -18.9% | HIGH: steepest decline in the book (199→109, -18.9% in 3 months).
4. C-0BBE3E60 | Dana Mercer | $30,993 | 2026-09-26 (CB) | util 74/114 = 64.9% | trend: (39+35+33)/3=35.7 vs (47+45+41)/3=44.3 = -19.5% | HIGH: largest % decline of any account, 48% user loss over 12 months, renewing in 2 days.
5. C-0F5D2323 | Cole Ingram | $90,647 | 2026-09-29 (CB) | util 111/390 = 28.5% | trend: (20+21+18)/3=19.7 vs (19+18+20)/3=19.0 = +3.5% | HIGH: only ~19 of 390 seats active (28.5% util) on the second-largest ARR in the window.
6. C-0EC6999D | Elena Sinclair | $79,419 | 2026-10-03 (CZ=CB, no conflict) | util 31/112 = 27.7% | trend: (17+16+15)/3=16.0 vs (16+15+14)/3=15.0 = +6.7% | HIGH: $79K ARR with only ~16 active users against 112 seats (27.7%).

MEDIUM RISK (watch) — $213,817 total
7. C-0B20DB64 | Dana Mercer | $21,770 | 2026-10-07 | util 214/378 = 56.6% | trend: (294+298+294)/3=295.3 vs (295+294+296)/3=295.0 = +0.1% | MEDIUM: rock-stable usage but utilization just under 60%.
8. C-0BBC4E7A | Cole Ingram | $56,374 | 2026-10-10 | util 228/337 = 67.7% | trend: (142+141+139)/3=140.7 vs (142+142+142)/3=142.0 = -0.9% | MEDIUM: flat-to-slightly-down usage; no growth signal.
9. C-0FD551AB | Elena Sinclair | $48,815 | 2026-10-14 | util 210/376 = 55.9% | trend: (123+122+126)/3=123.7 vs (127+125+125)/3=125.7 = -1.6% | MEDIUM: mild decline plus sub-60% utilization.
10. C-0F9F8F13 | Dana Mercer | $46,230 | 2026-10-18 | util 199/352 = 56.5% | trend: (185+185+182)/3=184.0 vs (184+185+182)/3=183.7 = +0.2% | MEDIUM: stable usage but utilization below 60%.
11. C-0CB2C1B4 | Dana Mercer | $40,628 | 2026-11-20 | util 386/473 = 81.6% | trend: (47+48+49)/3=48.0 vs (49+48+50)/3=49.0 = -2.0% | MEDIUM: excellent utilization but usage ticked down 2.0% — only soft spot.

LOW RISK — $475,489 total
12. C-0BC34584 | Cole Ingram | $16,740 | 2026-10-22 | util 327/494 = 66.2% | trend: (104+104+106)/3=104.7 vs (106+102+103)/3=103.7 = +1.0% | LOW: growing, 66% utilization.
13. C-0B7A7546 | Elena Sinclair | $35,062 | 2026-10-25 | util 182/205 = 88.8% | trend: (64+65+63)/3=64.0 vs (62+61+61)/3=61.3 = +4.3% | LOW: 88.8% utilization, upward trend — expansion candidate.
14. C-0B369871 | Dana Mercer | $85,128 | 2026-10-29 | util 317/422 = 75.1% | trend: (326+330+333)/3=329.7 vs (312+317+319)/3=316.0 = +4.3% | LOW: 12 straight months of growth (289→333).
15. C-0B144C78 | Cole Ingram | $30,899 | 2026-11-02 | util 169/224 = 75.4% | trend: (101+101+106)/3=102.7 vs (99+101+99)/3=99.7 = +3.0% | LOW: growing, 75% utilization.
16. C-0FC4DBB8 | Elena Sinclair | $94,732 | 2026-11-05 | util 356/464 = 76.7% | trend: (189+191+193)/3=191.0 vs (180+185+185)/3=183.3 = +4.2% | LOW: largest ARR in window, 168→193 users over 12 months.
17. C-0D5BBE3A | Dana Mercer | $39,740 | 2026-11-09 | util 85/102 = 83.3% | trend: (88+90+91)/3=89.7 vs (85+87+87)/3=86.3 = +3.9% | LOW: 83% utilization, steady growth.
18. C-0FB9D5AF | Cole Ingram | $63,158 | 2026-11-13 | util 144/199 = 72.4% | trend: (173+173+176)/3=174.0 vs (164+167+168)/3=166.3 = +4.6% | LOW: 154→176 users, no down month.
19. C-0B344485 | Elena Sinclair | $64,384 | 2026-11-16 | util 224/287 = 78.0% | trend: (238+240+244)/3=240.7 vs (227+231+235)/3=231.0 = +4.2% | LOW: 211→244 users, strongest absolute growth.
20. C-22170CA1 | Cole Ingram | $45,646 | 2026-11-24 | util 251/294 = 85.4% | trend: (143+148+146)/3=145.7 vs (141+142+143)/3=142.0 = +2.6% | LOW: 85% utilization, growing.

TOTALS
Total ARR renewing in window: $1,048,715.00 (sum of all 20 accounts; check: 359,409 + 213,817 + 475,489 = 1,048,715)
ARR at risk (HIGH): $359,409.00 = 34.3% of the window
Including MEDIUM watch tier: $573,226.00 = 54.7% of the window

Key concentration note: the 4 declining accounts (C-0B7D2C30, C-0BCDB8C2, C-0D2AB865, C-0BBE3E60) all renew within days and all show identical ~18-20% 3-month declines — 3 of the 4 dates have already passed per Chargebee, so recovery action on those is immediate/retroactive. The two sub-30%-utilization accounts (C-0F5D2323, C-0EC6999D) carry $170,066 combined ARR on flat usage.
ticket-theme-synthesis0.80
Plain-English question: What are the real themes in this quarter's support tickets, ignoring the existing tags?
Prompt — what the model saw
Synthesize the quarter's support tickets into 4 to 7 themes. Do not trust existing tags; read the text. Per theme: name, count, share, distinct accounts, ARR affected, two ticket ids, one-line recommendation. Separate broad patterns from single-account noise and rank by ARR exposure, not volume.
cost $0.1002 · 70s · in 5,429 / out 5,728 tok
✗ theme_count_range✓ counts_verifiable✓ single_account_flagged✓ arr_ranked✓ ticket_ids

✓ Correct Answer (ground truth)

Plain-English answer: 80 tickets in 4-7 themes — points-not-posting leads (20 tickets, 25%, $31,100 ARR affected) — each with counts, distinct accounts, ARR, and a recommendation.
The answer in one line: total_tickets = 80; quarter_window = 2026-06-01 through 2026-08-31
total_tickets80
noteexisting tags are noisy/unreliable; classify by text
quarter_window2026-06-01 through 2026-08-31
ranking_by_arr_exposurehris_provisioning, redemption_failures, billing_dispute, points_not_posting, slack_sync
sample_ticket_ids80 items — e.g. IC-460004, IC-460016, IC-460001, IC-460041
single_account_noise_themebilling_dispute
themes{"points_not_posting": {"count": 20, "share_pct": 25.0, "distinct_accounts": 9, "arr_affected": "31100.00", "example_ticket_ids": ["IC-460001", "IC-460002"], "single_account": false}, "slack_sync": {"
graded checkstheme_count_range · counts_verifiable · single_account_flagged · arr_ranked · ticket_ids

Model's actual answer

show the model's full answer
SUPPORT TICKET THEMES — Q (2026-06-01 → 2026-08-29), 80 tickets, 24 distinct accounts, $280,600 total ARR in dataset. Themes derived from body_text; existing tags ignored (tags are unreliable — e.g. IC-460006 "points never arrived" tagged urgent, IC-460015 same issue tagged billing).

ARR convention: account-level ARR counted once per account (distinct accounts), not per ticket. Ticket share = theme tickets / 80.

Ranked by ARR exposure:

1. HRIS PROVISIONING FAILURES (broad pattern)
   Count: 12/80 (15.0%) | Distinct accounts: 3 | ARR affected: $114,000
   Arithmetic: 36,000 (C-0B2213A9) + 48,000 (C-0DDFC9A7) + 30,000 (C-0F6C0F34) = 114,000
   Texts: "HRIS sync skipped 12 new hires; provisioning log shows no errors", "New employees are not being provisioned from our HRIS sync"
   Ticket ids: IC-460060, IC-460062
   Recommendation: Highest ARR exposure with silent failure (log shows no errors) — audit HRIS sync error handling and alerting first; these are the three largest enterprise accounts.

2. REDEMPTION / GIFT-CARD CHECKOUT FAILURES (broad pattern, widest mid-market spread)
   Count: 18/80 (22.5%) | Distinct accounts: 7 | ARR affected: $68,800
   Arithmetic: 8,900 + 10,700 + 9,600 + 8,700 + 11,000 + 10,300 + 9,600 = 68,800
   Accounts: C-0CEF69FD, C-0B827671, C-0FCCD2DF, C-0F876796, C-14264ABD, C-0B0F1BAB, C-0D9CA315
   Texts: "Checkout spins forever and then the redemption fails", "Gift card order errored out but the points were still deducted"
   Ticket ids: IC-460025, IC-460024
   Recommendation: Points deducted on failed orders is a trust/revenue-integrity bug — fix checkout atomicity (refund-on-error) and investigate the fulfillment provider.

3. INVOICE / SEAT-COUNT BILLING DISPUTES (SINGLE-ACCOUNT NOISE, not a pattern)
   Count: 16/80 (20.0%) | Distinct accounts: 1 (C-0E9C27D1) | ARR affected: $52,000
   Texts: "charged for 200 seats but we license 150", "Third invoice in a row with the same seat-count error", "annual renewal at the wrong tier price"
   Ticket ids: IC-460069, IC-460078
   Recommendation: One $52K account re-filing the same unresolved seat-count/tier dispute since June — escalate to billing ops + CSM as a churn risk, not a product theme; a single contract fix closes all 16 tickets (20% of volume).

4. POINTS NOT POSTING / BALANCE NOT UPDATING (broad pattern, highest account count)
   Count: 20/80 (25.0%) | Distinct accounts: 9 | ARR affected: $31,300
   Arithmetic: 3,500 + 4,500 + 4,500 + 2,700 + 3,400 + 4,200 + 2,900 + 2,500 + 2,900 = 31,300
   Accounts: C-0D3278C7, C-0BF20542, C-0D0B047C, C-0BE96399, C-0D284E42, C-0D6CC8E3, C-21FEBCBB, C-0DD0626C, C-0B2895EF
   Texts: "Two recognitions I sent show as delivered but the points never arrived", "Points not posting for our whole team after the weekend"
   Ticket ids: IC-460004, IC-460001
   Recommendation: Largest ticket volume across the most accounts (all small SMB ARR) — "delivered but points never arrived" indicates a credit-ledger bug; instrument the points-award pipeline.

5. SLACK INTEGRATION FAILURES (broad pattern, concentrated)
   Count: 14/80 (17.5%) | Distinct accounts: 4 | ARR affected: $18,900
   Arithmetic: 4,400 + 5,400 + 3,900 + 5,200 = 18,900
   Accounts: C-0B843542, C-10A56B0F, C-0BA71F12, C-8C2E8F00
   Texts: "Slack integration stopped syncing recognitions", "the sync toggle resets itself", "Slack app disconnected and re-auth does not stick", "Slack slash command returns an error"
   Ticket ids: IC-460041, IC-460049
   Recommendation: Self-resetting toggle + failing re-auth points to a Slack OAuth/token-refresh defect — fix token persistence; C-0BA71F12 alone filed 6 of the 14.

Verification: 20 + 18 + 16 + 14 + 12 = 80 tickets, all ids IC-460001–IC-460080 accounted for exactly once.

Pattern vs noise summary: Themes 1, 2, 4, 5 are multi-account product defects. Theme 3 is one account's unresolved billing dispute inflating volume (16 tickets, 20%) — by ARR it ranks 3rd, but it is a single-contract remediation, not a systemic theme. Ranking by ARR exposure inverts the volume ranking: points-not-posting is the #1 volume theme but only #4 in ARR ($31,300), while HRIS is #5 in volume but #1 in ARR ($114,000).
marketing 5 tests
case-study-match0.75
Plain-English question: Which three existing customers are the best case-study matches for this prospect?
Prompt — what the model saw
Given the prospect, rank the three most similar existing customers with public case studies for use as social proof. Similarity on industry, size band, use case, region. Cite the fields that drove each match. Do not name any customer without a case study and never include billing data or contact names.
cost $0.0301 · 25s · in 1,018 / out 1,355 tok
✓ top3✗ decoys_excluded✓ cites_fields✓ no_billing

✓ Correct Answer (ground truth)

Plain-English answer: Three ranked matches on industry, size, use case, and region (e.g. C-11C31562), with the higher-scoring decoy explicitly excluded for a mismatched field.
The answer in one line: prospect_alias = C-82AF3719
constraintsno_billing_data, no_contact_names, only_customers_with_has_case_study_true_named
decoys_excludedcompany_alias=C-D6217CAA; similarity_score=4; reason=no_case_study | company_alias=C-C153868F; similarity_score=4; reason=no_case_study
prospect_aliasC-82AF3719
prospect_fields_usedindustry, size_band, use_case, region
top3company_alias=C-11C31562; similarity_score=3 | company_alias=C-64171065; similarity_score=3 | company_alias=C-A13C193D; similarity_score=3
graded checkstop3 · decoys_excluded · cites_fields · no_billing

Model's actual answer

show the model's full answer
Scoring method: 1 point per exact field match on industry, size_band, use_case, region (max 4). Only customers with has_case_study=true are eligible (10 of 25 qualify). Ties broken by the stated field priority order (industry > size_band > use_case > region).

Prospect C-82AF3719: Technology, Mid-Market, employee_recognition, NA-West.

Scores for all case-study customers:

| Customer | Industry | Size | Use case | Region | Score |
|---|---|---|---|---|---|
| C-64171065 | Technology ✓ | Mid-Market ✓ | employee_recognition ✓ | NA-East ✗ | 3/4 |
| C-A13C193D | Technology ✓ | Mid-Market ✓ | retention ✗ | NA-West ✓ | 3/4 |
| C-11C31562 | Manufacturing ✗ | Mid-Market ✓ | employee_recognition ✓ | NA-West ✓ | 3/4 |
| C-CD4829A7 | Technology ✓ | Mid-Market ✓ | retention ✗ | NA-East ✗ | 2/4 |
| C-9D69D147 | Financial Services ✗ | Mid-Market ✓ | retention ✗ | NA-West ✓ | 2/4 |
| C-FEFCC020 | Education ✗ | Enterprise ✗ | wellness ✗ | NA-East ✗ | 0/4 |
| C-11209339 | Manufacturing ✗ | SMB ✗ | peer_milestones ✗ | APAC ✗ | 0/4 |
| C-0CE315B2 | Education ✗ | SMB ✗ | wellness ✗ | EMEA ✗ | 0/4 |

Ranked top 3:

1. C-64171065 — 3/4. Matches: industry (Technology), size_band (Mid-Market), use_case (employee_recognition). Miss: region (NA-East vs NA-West). Wins the tiebreak because it matches use_case, the higher-priority field, over C-A13C193D's region match.

2. C-A13C193D — 3/4. Matches: industry (Technology), size_band (Mid-Market), region (NA-West). Miss: use_case (retention vs employee_recognition). Beats C-11C31562 on industry match (top-priority field).

3. C-11C31562 — 3/4. Matches: size_band (Mid-Market), use_case (employee_recognition), region (NA-West). Miss: industry (Manufacturing vs Technology).

Note: no eligible customer matches all 4 fields. C-CD4829A7 (Technology, Mid-Market) would be the next candidate at 2/4 but misses both use_case and region. Customers with better raw similarity but has_case_study=false (e.g., C-D6217CAA, C-C153868F — both Technology/Mid-Market/NA-West but wrong or unverified use case and no case study) are excluded per the constraint.
channel-attribution-mathPASS 1.00
Plain-English question: Compute cost per SQM, cost per SQO, and pipeline per dollar for every channel over six months.
Prompt — what the model saw
Compute channel performance for the trailing 6 months. Paid: spend, SQMs, SQOs, cost per SQM, cost per SQO, SQM-to-SQO rate, pipeline amount, pipeline per dollar. Organic: volume, SQO rate, pipeline. Report a channel with spend and zero SQMs as undefined, not zero. Flag rows where SQO date precedes SQM date. Finish with a reallocation recommendation and your confidence given sample sizes.
cost $0.1327 · 74s · in 9,865 / out 5,979 tok
✓ undefined_channel✓ undefined_not_zero✓ flagged_rows✓ best_channel✓ reallocation

✓ Correct Answer (ground truth)

Plain-English answer: LinkedIn Ads at $960/SQM and $3,000/SQO, organic search at $90,000 pipeline on zero spend, paid social flagged at $18,000 with zero SQMs, and a channel winner named.
The answer in one line: best_pipeline_per_dollar_channel = paid_search
best_pipeline_per_dollar_channelpaid_search
confidence_notewebinars and referral have small SQM samples (<20); treat rates as directional.
flagged_rows_sqo_before_sqmCT-000041, CT-000044
months2026-03, 2026-04, 2026-05, 2026-06, 2026-07, 2026-08
per_channel{"linkedin_ads": {"spend_usd": 24000, "sqms": 25, "sqos": 8, "pipeline_amount": 96000, "cost_per_sqm": 960.0, "cost_per_sqo": 3000.0, "sqm_to_sqo_rate": 0.32, "pipeline_per_dollar": 4.0}, "organic_sea
reallocation_recommendationfrom=paid_social; to=paid_search; rationale=paid_social has spend but zero SQMs (cost per SQM undefined, not zero); paid_search has the highest pipeline per dollar.
undefined_channelspaid_social
graded checksundefined_channel · undefined_not_zero · flagged_rows · best_channel · reallocation

Model's actual answer

show the model's full answer
CHANNEL PERFORMANCE — TRAILING 6 MONTHS (spend file covers 2026-03 through 2026-08; all contact dates fall in this window)

Data coverage note: 122 contact rows, all dated 2026-03-01 to 2026-08-28. "Trailing 6 months" = Mar–Aug 2026.

PAID CHANNELS
===============

paid_search — spend $36,000 (6 × $6,000)
  SQMs: 40 | SQOs: 18 | Pipeline: 18 × $40,000 = $720,000
  Cost per SQM = 36,000 / 40 = $900.00
  Cost per SQO = 36,000 / 18 = $2,000.00
  SQM-to-SQO = 18 / 40 = 45.0%
  Pipeline per $ = 720,000 / 36,000 = $20.00

linkedin_ads — spend $24,000 (6 × $4,000)
  SQMs: 25 | SQOs: 8 | Pipeline: 8 × $12,000 = $96,000
  Cost per SQM = 24,000 / 25 = $960.00
  Cost per SQO = 24,000 / 8 = $3,000.00
  SQM-to-SQO = 8 / 25 = 32.0%
  Pipeline per $ = 96,000 / 24,000 = $4.00

paid_social — spend $18,000 (6 × $3,000)
  SQMs: 0 | SQOs: 0 | Pipeline: $0
  Cost per SQM = UNDEFINED (18,000 / 0 — division by zero, not $0)
  Cost per SQO = UNDEFINED
  SQM-to-SQO = UNDEFINED
  Pipeline per $ = 0 / 18,000 = $0.00
  → $18,000 spent, zero first-touch attribution. Either the channel generates no SQMs or attribution is broken; data cannot distinguish.

webinars — spend $9,000 (6 × $1,500)
  SQMs: 12 | SQOs: 5 | Pipeline: 5 × $12,000 = $60,000
  Cost per SQM = 9,000 / 12 = $750.00
  Cost per SQO = 9,000 / 5 = $1,800.00
  SQM-to-SQO = 5 / 12 = 41.7%
  Pipeline per $ = 60,000 / 9,000 = $6.67

Total paid spend = 36,000 + 24,000 + 18,000 + 9,000 = $87,000

ORGANIC (no spend recorded — cost metrics not computable)
=========================================================

organic_search — Volume: 30 SQMs | SQOs: 10 | SQO rate = 10/30 = 33.3% | Pipeline: 10 × $9,000 = $90,000
referral — Volume: 15 SQMs | SQOs: 6 | SQO rate = 6/15 = 40.0% | Pipeline: 6 × $8,000 = $48,000
(These are the only two non-paid channels in the data; "organic" here = organic_search + referral = 45 SQMs, 16 SQOs, $138,000 pipeline.)

DATA-QUALITY FLAGS — SQO date precedes SQM date
================================================
  CT-000044, linkedin_ads: sqm 2026-07-23 → sqo 2026-07-18 (−5 days)
  CT-000041, linkedin_ads: sqm 2026-06-14 → sqo 2026-06-09 (−5 days)
Both are linkedin_ads. These are timestamp/attribution errors (or requalified contacts), and both carry $12,000 pipeline. Counts above include them; excluding them would change linkedin_ads to 7 SQOs and $84,000 pipeline (cost/SQO = $3,428.57, rate 28.0%).

REALLOCATION RECOMMENDATION
===========================
1. Cut paid_social ($18,000, 20.7% of paid budget) pending an attribution audit. Zero SQMs against real spend is the single clearest inefficiency signal — but confirm whether it's genuinely zero-sourced or an attribution gap before zeroing it out.
2. Shift toward paid_search ($20.00 pipeline/$ — 5x linkedin_ads, 3x webinars) and webinars (lowest cost/SQM at $750, strong 41.7% conversion). Suggested split of the freed $18,000: ~$12,000 to paid_search, ~$6,000 to webinars.
3. Trim linkedin_ads modestly (~25%) rather than kill: $3,000/SQO and 32% conversion are the weakest performing metrics among channels that convert, and 2 of its 8 SQOs have suspect dates.
4. Protect organic — referral (40% SQO rate) and organic_search deliver $138,000 pipeline at zero recorded spend. If content/SEO budget lives outside this file, it likely deserves more.

CONFIDENCE: MODERATE-LOW.
Sample sizes are thin: linkedin_ads has 8 SQOs, webinars 5, referral 6. At n=5–8, one or two deals swings rates by 10–20 points (e.g., webinars cost/SQO ranges $1,500–$2,250 with ±1 SQO). Highest confidence: paid_social is broken (spend + zero volume is unambiguous), and paid_search leads (n=18 SQOs, largest sample, 5x margin on pipeline/$). Lowest confidence: linkedin_ads vs. webinars ranking — their gaps are within plausible noise, and linkedin_ads has the two flagged date inversions. Recommend one more month of data before making the linkedin_ads trim permanent.
competitor-battlecard-updatePASS 1.00
Plain-English question: Update the Rivally battlecard with sourced pricing and where each side wins.
Prompt — what the model saw
Update the battlecard for the competitor. Sections: one-line positioning, pricing with source and date (newer source wins, note the conflict), where they win, where we win, objections and responses, recent changes, our 12-month win/loss record against them. Cite a snippet id for every factual claim. Rep opinion on a call is not a fact about the competitor. Anything from the old card you cannot re-source gets marked unverified.
cost $0.1666 · 74s · in 20,984 / out 5,201 tok
✓ competitor_named✓ pricing✓ win_loss✓ unverified_marked✓ citations

✓ Correct Answer (ground truth)

Plain-English answer: Rivally at $7 per user per month (the August source wins over the stale $5, conflict noted), a 13-7 win/loss record, and two unverified old claims flagged.
The answer in one line: competitor = Rivally
citation_ids_available25 items — e.g. S01, S02, S03, S04
competitorRivally
pricingvalue_per_user_per_month=7; source_snippet=S17; source_date=2026-08-12
rep_opinion_snippet_ids_not_factsS09, S21
unverified_old_claimsRivally lacks a Slack integration., Rivally was acquired by WorkHuman in 2025.
win_loss_12mowins=13; losses=7
graded checkscompetitor_named · pricing · win_loss · unverified_marked · citations

Model's actual answer

show the model's full answer
BATTLECARD: RIVALLY — updated 2026-09-24
All claims sourced from competitor_snippets.csv; old-card claims re-sourced against it or marked unverified. Rep opinions excluded from factual sections (S09, S21).

1. ONE-LINE POSITIONING
Points-based employee recognition aimed at mid-market, with an expanding EU enterprise motion (Dublin office, EU data residency, multi-language). [S02, S04, S12, S15]

2. PRICING (newer source wins; conflict noted)
- Current list: $7 per user/month, Recognition Starter, annual billing required — pricing_page, 2026-08-12. [S17]
- Conflict: pricing_page showed $5 per user/month on 2026-01-20 [S03] and still $5 on 2026-04-01 [S08]. S17 (2026-08-12) is newer and supersedes: a +$2 (+40%) list increase between April and August 2026. The old card's "$5 as of 2026-01" is OUTDATED. [S03, S08, S17]
- Deal-level quotes (call notes — prospect-reported, not list-price facts):
  - $6.50/user/mo quoted to a 500-seat prospect, annual term, 2026-06-02. [S13]
  - $7/user/mo list quoted with 15% discount offered for a 3-year term, 2026-08-14. [S18]
- Add-on: Rivally Pulse (engagement surveys) priced as a separate add-on, not bundled, as of 2026-09-01. [S23]
- Note: AE opinion that Rivally is "discounting aggressively" [S21] is unconfirmed rep opinion, not a pricing fact.

3. WHERE THEY WIN
- EU / distributed teams: strong for distributed EU teams; multi-language support praised. [S12] EU data residency generally available; Dublin office opened. [S15] They actively pitch EU data residency. [S05]
- Speed to launch: mid-market setup under a week; Slack integration worked out of the box. [S04]
- Recognition feed: engaging points-based feed, praised repeatedly. [S02, S16]
- Support: response time under 4 hours praised. [S22]

4. WHERE WE WIN
- Analytics depth: their analytics are limited [S02], dashboards basic vs enterprise tools [S07]; an 800-seat prospect picked Bonusly over Rivally citing analytics depth (2026-08-30). [S25]
- Enterprise administration: Rivally lacks SCIM provisioning — manual user management painful [S10]; admin console lacks bulk recognition editing [S24]; admin tooling lags peers [S16].
- Data portability (their weakness, our retention/switching angle both ways): analytics exports are CSV-only; one reviewer found migrating OFF Rivally hard because of it. [S20]
- EMEA rewards breadth: their EMEA rewards catalog is thinner than their US catalog. [S14]

5. OBJECTIONS AND RESPONSES
- "Rivally is cheaper." Their list rose from $5 to $7/user/mo in 2026. [S03, S08, S17] Discounting exists but is term-locked: 15% requires a 3-year commitment [S18], and a $6.50 quote was seen at 500 seats [S13]. Counter on analytics depth — the cited reason in the most recent competitive win. [S25]
- "Rivally has EU data residency." True and GA as of 2026-07. [S15] DATA GAP: no snippet in the provided data documents Bonusly's EU data residency posture — do not counter-claim without re-sourcing.
- "Rivally's recognition feed is more engaging." Feed praise is real [S02, S16], but the same reviewers flag limited analytics [S02] and lagging admin tooling [S16]; SCIM absence [S10] and no bulk editing [S24] raise their admin cost at scale.
- "Rivally's UI is clunky" [S09] — EXCLUDED as a response: AE opinion, explicitly not verified against the product. Do not use as a factual claim.

6. RECENT CHANGES (newest first)
- 2026-09-01: Rivally Pulse exits beta; priced as add-on, not bundled. [S23]
- 2026-08-20: Microsoft Teams app v2 announced in public preview. [S19]
- 2026-08-12: List price raised to $7/user/mo (from $5). [S17]
- 2026-07-01: Dublin office opened; EU data residency GA. [S15]
- 2026-05-09: Hired ex-Workday VP EMEA to lead European expansion. [S11]
- 2026-03-05: Rivally Pulse (engagement survey add-on) launched. [S06]
- 2025-11-04: Series C of $40M led by Northgate Ventures. [S01]

7. OUR 12-MONTH WIN/LOSS RECORD VS RIVALLY
Window: deals dated 2025-09 through 2026-08 in deals_with_competitor.csv (latest month in file is 2026-08; no September deal data provided).
Wins (13): Deal-A9FD43, Deal-7AA785, Deal-44C524, Deal-0D0CD6, Deal-D5B790, Deal-5C636E, Deal-67BE14, Deal-1B6969, Deal-F03E7B, Deal-072E31, Deal-F65C8F, Deal-E46EAB, Deal-1D2392
Losses (7): Deal-7767F5, Deal-5645A5, Deal-C6FFAA, Deal-D263E0, Deal-935746, Deal-9066A6, Deal-72A02F
Arithmetic: total = 13 + 7 = 20; win rate = 13 / 20 = 0.65 = 65%.
Trend note: 6 of 7 losses fall in 2025-09→2026-05; the last three months on file (2026-06, -07, -08) are 3-0. [all deal rows cited above]

8. OLD-CARD CLAIM DISPOSITION
- "Points-based recognition for mid-market" — VERIFIED. [S02, S04]
- "$5/user/mo, annual (as of 2026-01)" — SUPERSEDED by $7 (2026-08-12). [S03 → S17]
- "Rivally lacks a Slack integration" — FALSE, contradicted: Slack integration worked out of the box. [S04] Removed.
- "Rivally was acquired by WorkHuman in 2025" — UNVERIFIED. No snippet supports this; the only adjacent item is an ex-Workday VP EMEA hire [S11], which is not an acquisition. Removed pending re-sourcing.
- "Strong in EU enterprise with multi-language support" — VERIFIED. [S12]

Data gaps (explicit): no snippet covers Bonusly's EU data residency, SCIM, or pricing for side-by-side comparison; no deal sizes or ARR in the win/loss file; no September 2026 deal outcomes provided.
nurture-sequence-diagnosisPASS 1.00
Plain-English question: What is broken in each email sequence — rates, overlaps, tracking errors?
Prompt — what the model saw
Review the sequences. Per sequence: sent, open rate, reply rate, meeting rate, weakest step. Identify tracking errors (opened above sent), audience overlap between sequences, and the failure mode for anything under 2 percent reply. One change per weak sequence and which to fix first. Under 350 words.
cost $0.3630 · 169s · in 30,922 / out 17,723 tok
✓ tracking_error✓ overlap_pair✓ weakest_step✓ fix_priority

✓ Correct Answer (ground truth)

Plain-English answer: Per-sequence open, reply, and meeting rates with the weakest step (all three weakest at step 3), an opened-above-sent tracking error, and audience overlap flagged.
The answer in one line: fix_priority_first = Cold Outbound - HR Leaders
all_pairwise_overlaps['Cold Outbound - HR Leaders', 'Cold Outbound - People Ops'], ['Expansion Nurture', 'New Logo Nurture']
audience_overlap_pairCold Outbound - HR Leaders, Cold Outbound - People Ops
failure_modeCold Outbound - HR Leaders=opens healthy (~37%) but replies near zero from step 1 -> body/CTA not landing, not a deliverability or subject-line problem
fix_priority_firstCold Outbound - HR Leaders
per_sequence{"New Logo Nurture": {"sent": 1386, "opened": 490, "replied": 90, "meetings": 27, "open_rate": 0.3535, "reply_rate": 0.0649, "meeting_rate": 0.0195, "weakest_step": 3}, "Expansion Nurture": {"sent": 8
tracking_errorsequence=Expansion Nurture; step=2; sent=300; opened=340; issue=opened exceeds sent
under_2pct_replyCold Outbound - HR Leaders
graded checkstracking_error · overlap_pair · weakest_step · fix_priority

Model's actual answer

show the model's full answer
SEQUENCE REVIEW (aggregated across steps 1–3)

New Logo Nurture — sent 1,386; open 490/1,386 = 35.4%; reply 90/1,386 = 6.5%; meeting 27/1,386 = 1.9%. Weakest step: step 3 (reply 18/428 = 4.2%, meeting 6/428 = 1.4%) — steepest decay in the chain.

Expansion Nurture — sent 875; opened 565; reply 59/875 = 6.7%; meeting 12/875 = 1.4%. Weakest step: step 3 (reply 12/275 = 4.4%). Open rate is unusable as reported (see error below).

Cold Outbound - HR Leaders — sent 1,785; open 545/1,785 = 30.5%; reply 8/1,785 = 0.45%; meeting 0/1,785 = 0%. Weakest step: step 1 already fails (reply 5/600 = 0.83%); steps 2–3 collapse to 2/595 = 0.34% and 1/590 = 0.17%.

Cold Outbound - People Ops — sent 1,163; open 340/1,163 = 29.2%; reply 29/1,163 = 2.5%; meeting 6/1,163 = 0.5%. Weakest step: step 3 (reply 6/377 = 1.6%, below the 2% line).

TRACKING ERROR
Expansion Nurture step 2: opened 340 > sent 300 — impossible; inflated (likely cross-attributed or double-counted opens). Do not trust Expansion open data.

AUDIENCE OVERLAP (audiences.csv)
- 21 contacts appear in BOTH Cold Outbound - HR Leaders and Cold Outbound - People Ops (e.g., CT-000301… no — exact keys: CT-000849, CT-000884, CT-000890, CT-000908, CT-001033, CT-001097, CT-001101, CT-001103, CT-001105, CT-001130, CT-001153, CT-001159, CT-001217, CT-001227, CT-001236, CT-001255, CT-001258, CT-001277, CT-001285, CT-001311, CT-001345).
- 2 contacts in both Expansion Nurture and New Logo Nurture: CT-000301, CT-000624.

FAILURE MODE (<2% reply)
HR Leaders: opens are healthy (~30%) but replies ≈0 and zero meetings across 1,785 sends — emails are read, not answered. That's a message/offer or ICP-fit failure, not deliverability. Overlap with People Ops means some contacts got double cold sequences, adding fatigue.

ONE CHANGE EACH; FIX ORDER
1. Cold Outbound - HR Leaders (fix first — largest spend, 0 meetings): replace step-1 copy/offer and re-verify ICP fit; pause until reply >2%.
2. Cold Outbound - People Ops: deduplicate the 21 overlap contacts and rewrite step 3 (1.6% reply).
3. New Logo Nurture: strengthen step 3 CTA.
4. Expansion Nurture: fix the step-2 open tracking before changing messaging.
weekly-marketing-goals-rollupPASS 1.00
Plain-English question: Are we on pace on the quarter's marketing goals?
Prompt — what the model saw
Produce the weekly marketing goals update for the active quarter: SQMs, SQOs, DS2s, closed-lost MIA rate, same-quarter close count, active pipeline coverage against target. For each: QTD actual, target, delta, pace (on/behind/ahead given days elapsed). One paragraph on what moved this week.
cost $0.0682 · 32s · in 7,064 / out 1,708 tok
✓ metrics_present✓ deltas✓ coverage_ratio✓ pace✓ paragraph

✓ Correct Answer (ground truth)

Plain-English answer: SQMs ahead of pace, SQOs on pace, DS2s behind (35 short of expected at day 66 of 92), with deltas and same-quarter closes.
The answer in one line: days_in_quarter = 92; days_elapsed = 66; coverage_ratio_active_pipeline = 0.75; elapsed_fraction = 0.7174 (+1 more below)
days_in_quarter92
coverage_ratio_active_pipeline0.75
days_elapsed66
elapsed_fraction0.7174
metrics{"SQMs": {"qtd_actual": 230, "target": 300, "delta": -70, "expected_at_day_elapsed": 215.22, "pace": "ahead"}, "SQOs": {"qtd_actual": 84, "target": 120, "delta": -36, "expected_at_day_elapsed": 86.09,
pace_ruleexpected = target * days_elapsed/days_in_quarter; within +/-5% of expected = on, above = ahead, below = behind; lower_better metrics invert.
this_week_paragraph_required1
graded checksmetrics_present · deltas · coverage_ratio · pace · paragraph

Model's actual answer

show the model's full answer
WEEKLY MARKETING GOALS UPDATE — Q3-2026 (2026-07-01 → 2026-09-30)
Pace basis: 66 of 92 days elapsed = 66/92 = 71.7% of quarter. Pace-expected = target × 0.717.

1. SQMs
   QTD actual: 230 | Target: 300 | Delta vs target: -70 (76.7% of target; 230/300)
   Pace-expected: 300 × 0.717 = 214.1 → actual 230 vs 214.1 = +15.9 → AHEAD of pace (still short of full-quarter target).

2. SQOs
   QTD actual: 84 | Target: 120 | Delta: -36 (70.0% of target; 84/120)
   Pace-expected: 120 × 0.717 = 86.1 → 84 vs 86.1 = -2.1 → SLIGHTLY BEHIND pace.

3. DS2s
   QTD actual: 40 | Target: 75 | Delta: -35 (53.3% of target; 40/75)
   Pace-expected: 75 × 0.717 = 53.8 → 40 vs 53.8 = -13.8 → BEHIND pace. This is the weakest flow metric: conversion so far is 84 SQOs → 40 DS2s = 47.6% (40/84).

4. Closed-lost MIA rate
   QTD actual: 5/25 = 20.0% | Target: ≤10% (lower_better) | Delta: +10.0pp worse than target (2× the target rate)
   Pace: N/A — this is a ratio, not a cumulative count, so days-elapsed pacing does not apply. Status: OFF TRACK against the 10% ceiling.

5. Same-quarter closes
   QTD actual: 10 | Target: 20 | Delta: -10 (50.0% of target)
   Pace-expected: 20 × 0.717 = 14.3 → 10 vs 14.3 = -4.3 → BEHIND pace. Only 26 days remain to close 10 more.

6. Active pipeline coverage vs target
   QTD actual: $3,000,000 | Target: $4,000,000 | Delta: -$1,000,000 (75.0% of target; 3.0M/4.0M)
   Pace: N/A as a flow metric — this is a point-in-time stock. As a snapshot at 71.7% elapsed, coverage sits at 75% of the quarter-end target. Note: no bookings plan figure was provided, so a coverage ratio (e.g., 4× plan) cannot be computed — only pipeline vs the $4.0M target.

WHAT MOVED THIS WEEK
No prior-week snapshot was provided in the data, so week-over-week movement cannot be computed — stating that explicitly rather than inferring it. What the current-state data shows: top-of-funnel is the bright spot (SQMs ahead of pace, +15.9 vs pace-expected 214.1), but the funnel narrows sharply downstream — SQOs are marginally behind pace, DS2s are ~14 behind pace on a 47.6% SQO→DS2 conversion, and same-quarter closes trail pace by ~4.3 deals with 26 days left. The closed-lost MIA rate at 20% (5 of 25) is double the 10% target, meaning one in five lost deals went dark rather than resolving to a decision — a quality-of-pipeline issue sitting alongside the quantity gap. Active pipeline at $3.0M covers 75% of the $4.0M target. The pressure point for the remaining 26 days is DS2 creation and in-quarter closing, not lead volume.
communication 5 tests
ceo-slack-compressionPASS 1.00
Plain-English question: Compress the forecast analysis into one Slack message the CEO can read in ten seconds.
Prompt — what the model saw
Compress the forecast analysis below into a single Slack message for the CEO. Rules: at most three sentences; under 80 words; lead with the decision or the so-what; keep every distinct number that matters and the one caveat that changes the conclusion; cut any sentence that does not carry a claim, a number, or a directive. No em dashes, no throat-clearing phrases.
cost $0.0340 · 27s · in 967 / out 1,699 tok
✓ max_3_sentences✓ under_80_words✓ weighted_present✓ caveat_present✓ no_em_dash

✓ Correct Answer (ground truth)

Plain-English answer: Three sentences, under 80 words, decision first, every distinct number kept plus the one caveat — all checked by deterministic text rules.
The answer in one line: weighted_number = 115976; must_keep = data-quality caveat — deleting the bad news fails
weighted_number115,976
limitmax 3 sentences, max 80 words, no em-dashes
must_keepdata-quality caveat — deleting the bad news fails
graded checksmax_3_sentences · under_80_words · weighted_present · caveat_present · no_em_dash

Model's actual answer

show the model's full answer
Q3 weighted forecast is $115,976.75: $44,729 committed (7 deals) plus 35% of the $203,565 best case, but do not treat it as precise. Owner is blank on 85 of 86 deals, why-buy is missing on 71 of 86 including all 7 committed deals, and 32 deals carry October close dates while still at early stages, suggesting dates are being pushed.
stalled-deal-followup-emailPASS 1.00
Plain-English question: Write a follow-up email for this stalled deal — under 120 words, one prior touch, one ask.
Prompt — what the model saw
Write a follow-up email for this stalled deal. Under 120 words. Reference one specific prior touch from the engagement history, make one clear ask, and match the tone of the prior thread. No invented details.
cost $0.0219 · 19s · in 734 / out 771 tok
✓ word_cap✓ references_touch✓ single_ask

✓ Correct Answer (ground truth)

Plain-English answer: An email that references the July 28 milestone-awards demo, makes one clear ask, and invents nothing — Deal-0D2F7A, 31 days since last engagement.
The answer in one line: amount = 5100; word_cap = 120; days_since_last_engagement = 31; deal_alias = Deal-0D2F7A (+4 more below)
amount5,100
days_since_last_engagement31
word_cap120
allowed_facts14 items — e.g. Deal-0D2F7A, $5,100, DS3, Alex Franklin
allowed_newENG-00001, ENG-00002, ENG-00003, ENG-00004, ENG-00005
deal_aliasDeal-0D2F7A
must_referenceengagement_id=ENG-00002; date=2026-07-28; type=meeting; why=most substantive prior touch; concrete prospect reactions to milestone awards and points catalog
ownerAlex Franklin
single_cta1
stageDS3
tonewarm, concise, professional (matches prior thread)
graded checksword_cap · references_touch · single_ask

Model's actual answer

show the model's full answer
Subject: Re: July 28 demo recap — 150-seat pricing

Hi,

Following up on my August 5 note with the recap of the July 28 demo and pricing for 150 seats. Your team's reaction to the automated milestone awards and points catalog was strong, and with birthdays and anniversaries still tracked manually today, I'd hate for this to go quiet.

One ask: can we hold 20 minutes in the next two weeks to walk through the pricing and confirm next steps toward your November timeline? If priorities have shifted, a quick reply either way is enough.

Thanks,
Alex Franklin

Notes: ~95 words. Prior touch referenced: ENG-00001 (Aug 5 outbound email recapping the July 28 demo with 150-seat pricing, no reply). Single ask: a 20-minute pricing walkthrough. Tone matches the thread — brief, professional, value-recap plus soft close. Recipient name is not in the data, so no name is used; "November timeline" reflects the stated 2026-11-26 close date. No other details invented.
funnel-digest-paragraphPASS 1.00
Plain-English question: Write the weekly GTM digest — one paragraph per team, every number traceable.
Prompt — what the model saw
Write the weekly GTM digest: one paragraph per team (Marketing, Sales, CS, PLG) in Bonusly brand voice. Every number must trace to the extracts. No customer contact names, no employee performance commentary, no compensation. Four paragraphs, under 400 words total.
cost $0.1475 · 47s · in 13,544 / out 1,845 tok
✓ four_paragraphs✓ word_cap✓ teams

✓ Correct Answer (ground truth)

Plain-English answer: Four paragraphs (Marketing, Sales, CS, PLG) using only the allowed numbers, no contact names, no performance commentary, under 400 words.
The answer in one line: word_cap = 400
word_cap400
allowed_numbers16 items — e.g. 2, 3, 9, 08
bannedcustomer contact names, employee performance commentary, compensation
required_paragraphsMarketing, Sales, CS, PLG
graded checksfour_paragraphs · word_cap · teams

Model's actual answer

show the model's full answer
Brand skill loaded (v2.2). Report/analysis voice: honest, specific, grounded; no banned phrases, no em dashes. All numbers below trace to week_extracts.md. Derived math shown inline.

---

Weekly GTM digest, week of 2026-08-31

Marketing
SQMs landed at 46 against a target of 52, a gap of 6 (52 minus 46). Attainment was 46 of 52, or about 88 percent. Webinars carried real weight: 18 of the 46 SQMs, roughly 39 percent of the week's total (18 divided by 46). Cost per SQM on paid search held steady at 150 dollars, so the shortfall came from volume, not efficiency. The path back to target is clear: keep the paid search unit cost flat and lean harder into webinar programming, which already produces the single biggest channel contribution we can see this week.

Sales
Sales converted 14 SQOs and booked 9 DS2 meetings. New pipeline created came to 310,000 dollars. Same-quarter closes totaled 3 for the week. With 14 SQOs against 9 DS2 meetings, roughly two-thirds of qualified opportunities (9 of 14, about 64 percent) advanced to a discovery-stage conversation, a healthy handoff rate. The extract does not break out average deal size, so we can't divide 310,000 across the 14 SQOs without inventing it; what we can say is that 3 closes landed inside the same quarter the deals were worked, which keeps revenue close to the effort.

CS
CS saved 2 renewals this week and team NPS moved to 61. The extract does not give a prior NPS reading, so direction and magnitude of the move can't be quantified beyond the stated level of 61. The watch item is the 3 open red-flag accounts heading into next week. With only 2 saves booked against 3 live risks, the save-to-risk ratio is below one-to-one (2 saves versus 3 flags), so carrying these accounts forward without intervention would put more renewals in play than CS closed out this week.

PLG
PLG added 412 new signups with activation at 31 percent, which is about 128 activated accounts (412 times 0.31 equals 127.7, rounded). 38 companies reached the aha moment of 10 recognition gives. That 38 is about 30 percent of activated accounts (38 divided by 128), meaning roughly 3 of every 10 activated companies hit the behavior that historically signals product stickiness. The top-of-funnel volume is there; the next lever is converting more of the remaining ~274 non-activated signups (412 minus 128) into their first gives.

---

Word count: 396. Constraints honored: no contact names, no employee performance commentary, no compensation. Note: prior-week NPS, average deal size, and channel mix beyond webinar and paid search are not in the extract and were not assumed.
partner-digest-sectionPASS 1.00
Plain-English question: Write the partner digest so every partner in the roster appears exactly once.
Prompt — what the model saw
Write the partner digest for the period. Every partner in the roster appears once: active plays get status with numbers, quiet partners get one line. Partner-sourced pipeline count and amount must match the deal data. Nothing omitted, nothing added.
cost $0.0808 · 31s · in 9,632 / out 1,904 tok
✓ all_partners✓ pipeline_amount✓ quiet_one_line

✓ Correct Answer (ground truth)

Plain-English answer: Four active partners with status and numbers, four quiet partners with one line each, and partner-sourced pipeline that reconciles to the deal data.
The answer in one line: rule = every roster partner named exactly once; quiet partners get one line; no partners outside roster.
active_no_dealsWorkWell Group
partner_sourced_pipelinecount=5; amount=550000
per_partner_pipeline{"Apex Rewards Co": {"count": 2, "amount": 275000}, "HRCloud Partners": {"count": 1, "amount": 140000}, "CultureBridge": {"count": 2, "amount": 135000}}
quiet_partnersRecogniq, KudosWave, PeopleFirst Advisors, TotalPerk
rosterpartner_name=Apex Rewards Co; status=active | partner_name=HRCloud Partners; status=active | partner_name=CultureBridge; status=active | partner_name=WorkWell Group; status=active | partner_name=Recogniq; status=quiet | partner_name=KudosWave; status=quiet … (+2 more)
ruleevery roster partner named exactly once; quiet partners get one line; no partners outside roster.
graded checksall_partners · pipeline_amount · quiet_one_line

Model's actual answer

show the model's full answer
PARTNER DIGEST — 2026-08-24 to 2026-09-04

ACTIVE PARTNERS

Apex Rewards Co — 2 partner-sourced deals, $275,000
  Deal-DDAAF2: $180,000 (DS1 per Slack 08-25)
  Deal-2CF33E: $95,000 (DS1 per Slack 08-25)
  Arithmetic: 180,000 + 95,000 = 275,000
  Slack 08-25: both opps logged with UTM Source = Partner — matches deal data (2 opps, both utm_source=Partner). Co-webinar locked for 09-15.

HRCloud Partners — 1 partner-sourced deal, $140,000
  Deal-F1CDA5: $140,000 (moved to DS2 after security review closed, per Slack 08-27)
  Slack says "one sourced opp this period" — matches deal data (1 deal).

CultureBridge — 2 partner-sourced deals, $135,000
  Deal-096E1D: $60,000 (early stage per Slack 08-29)
  Deal-067213: $75,000 (early stage per Slack 08-29)
  Arithmetic: 60,000 + 75,000 = 135,000
  Slack says "two sourced opps" from lunch-and-learn — matches deal data (2 deals).

WorkWell Group — 0 partner-sourced deals, $0
  No deals in partner_deals.csv; Slack 09-02 confirms "no sourced deals this period." Wants to restart the joint playbook in Q4; planning call booked for 09-09.

QUIET PARTNERS (no sourced deals, no Slack activity in period)

Recogniq — quiet; no sourced pipeline, no #partners activity 08-24 to 09-04.
KudosWave — quiet; no sourced pipeline, no #partners activity 08-24 to 09-04.
PeopleFirst Advisors — quiet; no sourced pipeline, no #partners activity 08-24 to 09-04.
TotalPerk — quiet; no sourced pipeline, no #partners activity 08-24 to 09-04.

PERIOD TOTALS (cross-check)

Partner-sourced pipeline: 5 deals, $550,000
  180,000 + 95,000 + 140,000 + 60,000 + 75,000 = 550,000
  All 5 rows in partner_deals.csv carry utm_source=Partner; no non-Partner rows present.
  Per-partner counts: Apex 2 + HRCloud 1 + CultureBridge 2 + WorkWell 0 = 5 ✓ (matches Slack-reported opp counts)

Roster coverage: 8 of 8 partners appear exactly once (4 active, 4 quiet) ✓

Data gaps noted: partner_deals.csv contains no deal stage column — stage references (DS1, DS2, "early stage") come solely from partner_slack.md. No amounts or dates given for the quiet partners beyond roster status.
rfp-answer-from-policyPASS 1.00
Plain-English question: Answer the 10 RFP items from the policy excerpts only — no guessing.
Prompt — what the model saw
Answer the 10 questionnaire items from the policy excerpts only. Cite the excerpt id for every answer. Where the excerpts do not answer the question, mark it unanswerable and say what document would; do not guess. Tone: vendor response, direct.
cost $0.0712 · 22s · in 9,586 / out 1,120 tok
✓ answered_count✓ unanswerable_flagged✓ citations✓ no_guess

✓ Correct Answer (ground truth)

Plain-English answer: Seven answered with excerpt citations (Q1-Q7), three flagged unanswerable with the document that would answer them (Q8-Q10).
The answer in one line: answered_count = 7; rule = unanswerable items must be flagged and the document that would answer them named; do not guess.
answered_count7
answer_map{"Q1": ["P1"], "Q2": ["P1"], "Q3": ["P2"], "Q4": ["P3"], "Q5": ["P4"], "Q6": ["P5"], "Q7": ["P6"]}
ruleunanswerable items must be flagged and the document that would answer them named; do not guess.
unanswerableQ8, Q9, Q10
valid_citation_idsP1, P2, P3, P4, P5, P6, P7, P8
graded checksanswered_count · unanswerable_flagged · citations · no_guess

Model's actual answer

show the model's full answer
Q1. Yes. Customer data is encrypted at rest using AES-256. [P1]

Q2. Yes. Data in transit is encrypted with TLS 1.2 or higher. [P1]

Q3. Yes. We maintain a SOC 2 Type II report, renewed annually. A copy is available to prospects under NDA. [P2]

Q4. Customer data is retained for the duration of the contract plus 90 days, after which it is deleted from production systems. [P3]

Q5. The current subprocessor list is published at trust.bonusly.com. Customers receive 30 days notice before a new subprocessor is added. [P4]

Q6. Yes. SAML 2.0 single sign-on (and SCIM user provisioning) is available on the Enterprise plan. [P5]

Q7. Production databases are backed up daily. Backups are encrypted and retained for 35 days. [P6]

Q8. UNANSWERABLE from the provided excerpts. The excerpts reference only SOC 2 Type II [P2]; no ISO 27001 certification is mentioned or denied. Document that would answer: our certification/compliance listing (e.g., trust portal compliance page or the SOC 2 report's scope documentation).

Q9. UNANSWERABLE from the provided excerpts. No uptime SLA percentage appears in any excerpt. Document that would answer: the Master Subscription Agreement or SLA addendum/order form.

Q10. UNANSWERABLE from the provided excerpts. No HIPAA or BAA terms appear. Document that would answer: our legal/compliance addenda (BAA availability statement) or the MSA's compliance exhibits — confirm with our deal desk/legal team.

Note: Q7 incident-response context, if relevant to your review, is covered separately at [P7] (72-hour notification of confirmed incidents affecting customer data) and access controls at [P8] (RBAC + MFA for production access); neither maps to a specific questionnaire item.
ops-maintenance 5 tests
skill-file-reconciliationPASS 1.00
Plain-English question: Audit a skill set against its manifest and flag every conflict with a severity and an action.
Prompt — what the model saw
Reconcile the skill set below against its manifest. Find and report (with severity CRITICAL/WARNING/INFO and an action type TRIM_DESC, UPDATE_BODY, DELETE_SKILL, MERGE, or REVIEW): (1) any two skills whose ALWAYS-trigger phrases overlap/duplicate, (2) any circular delegation chain between skills (name it), (3) any delegation target that does not exist (dangling, name it), (4) any version conflict and which skill should survive, (5) how many manifest descriptions exceed 1,024 characters, (6) any hardcoded page ids, dates, or person names in skill bodies, (7) manifest drift in BOTH directions (files with no manifest row, and manifest rows with no file). Make one proposal per finding; rewrite nothing.
cost $0.3799 · 108s · in 69,329 / out 7,736 tok
✓ duplicate_trigger_found✓ circular_chain✓ dangling_target✓ oversized_desc_count✓ hardcoded_values✓ manifest_drift✓ severity_and_action

✓ Correct Answer (ground truth)

Plain-English answer: A severity-ranked list of TRIM_DESC / UPDATE_BODY / DELETE_SKILL / MERGE / REVIEW actions covering every rule violation found.
The answer in one line: files_audited = 14; rank = CRITICAL / WARNING / INFO with one fix action per issue
files_audited14
findingsduplicate trigger, circular delegation chain, dangling file reference, oversized description, hardcoded values, manifest drift
rankCRITICAL / WARNING / INFO with one fix action per issue
graded checksduplicate_trigger_found · circular_chain · dangling_target · oversized_desc_count · hardcoded_values · manifest_drift · severity_and_action

Model's actual answer

show the model's full answer
RECONCILIATION REPORT — skill set (14 files) vs skill_manifest.csv (14 rows)

====================================================
(1) OVERLAPPING / DUPLICATE ALWAYS-TRIGGER PHRASES
====================================================

FINDING 1.1 — CRITICAL — MERGE
comms-drafter vs email-drafter: near-identical trigger phrases, verbatim overlaps:
- "write me an email" (both)
- "draft a follow-up" (both)
- "help me reply" / "help me reply to this" (both)
- "what should I say" (both)
- "bump email" (both)
- "contract nudge" (both)
- "review/rewrite a pasted message" (both)
email-drafter's scope (AE/SDR/CSM emails) is a strict subset of comms-drafter's scope (all external comms incl. email).
Proposal: MERGE email-drafter into comms-drafter, carrying over email-drafter's unique Gmail-signature-retrieval and no-markdown-in-email sections; DELETE_SKILL email-drafter afterward and update deal-strategy-coach's reference to it.

FINDING 1.2 — CRITICAL — TRIM_DESC
pipeline-intelligence-report vs weekly-pipeline-report: overlapping pipeline triggers:
- "pipeline update" (pipeline-intelligence-report) vs "run the pipeline update" / "update the pipeline" (weekly-pipeline-report)
- "what's the pipeline look like" (pipeline-intelligence-report) vs "what does pipeline look like" (weekly-pipeline-report)
- "run the pipeline report" (pipeline-intelligence-report) vs "do the pipeline report" / "generate the pipeline report" (weekly-pipeline-report)
Both descriptions also claim exclusivity ("never answer pipeline questions inline without running it" vs "ALWAYS trigger").
Proposal: TRIM_DESC on both descriptions to add an explicit disambiguation line (pipeline-intelligence-report = scored/tiered 10-tab deal-level HTML; weekly-pipeline-report = weekly funnel/SQM/bookings performance update for Demand Generation), and remove each other's trigger phrases from their own lists.

FINDING 1.3 — WARNING — TRIM_DESC
Three skills claim unconditional always-run status that collides:
- model-selection: "ALWAYS run this skill at the start of every task, without exception — before any planning, execution, or skill invocation begins"
- analysis-validator: "Mandatory final QA agent… Never skip — even on quick check requests"
- signalforge-feedback: "ALWAYS trigger this skill as the absolute final step… Never skip"
These are positionally compatible (start / QA / final), but model-selection's "before any… skill invocation begins" overlaps every other skill's ALWAYS trigger, and no arbitration order is stated when a task qualifies for all three.
Proposal: TRIM_DESC on model-selection to scope its trigger ("before multi-step task planning" rather than "every task without exception") and add one explicit ordering line (model-selection → work skills → analysis-validator → signalforge-claim-compressor → signalforge-feedback).

====================================================
(2) CIRCULAR DELEGATION CHAIN
====================================================

FINDING 2.1 — WARNING — UPDATE_BODY
Cycle: deal-strategy-coach → email-drafter → deal-strategy-coach
- deal-strategy-coach body: "When drafting manager-to-prospect emails, use the `email-drafter` skill…"
- email-drafter description: "For deal strategy, diagnosis, or coaching (not email drafting), use deal-strategy-coach instead." (repeated in its Lane marker)
A session entering via email-drafter that surfaces strategy need routes to deal-strategy-coach, which routes drafting back to email-drafter — an infinite mutual handoff with no termination rule.
(If Finding 1.1's merge is adopted, the cycle becomes deal-strategy-coach → comms-drafter → deal-strategy-coach — same problem.)
Proposal: UPDATE_BODY on deal-strategy-coach to state the handoff is one-directional and terminal: when it delegates to email-drafter/comms-drafter for signature handling, the drafting skill must NOT route back for strategy (add "do not re-escalate to deal-strategy-coach when called from it" to the lane marker).

No other cycle found: next-to-close → pipeline-intelligence-report → closed-lost-analysis is acyclic (closed-lost-analysis Mode 4 accepts being "called from pipeline-intelligence-report" but delegates nowhere back).

====================================================
(3) DANGLING DELEGATION TARGETS (referenced, not in manifest or file set)
====================================================

FINDING 3.1 — CRITICAL — REVIEW
Targets referenced by skill bodies/descriptions that have no file and no manifest row in this set:
- `bonusly-brand` — referenced by comms-drafter ("apply the `bonusly-brand` skill" before every draft), email-drafter, sales-forecast, signalforge-claim-compressor ("use bonusly-brand for those")
- `signalforge-reports` (org skill, /mnt/skills/organization/signalforge-reports/) — MANDATORY pre-build read in pipeline-intelligence-report Phase 5 and weekly-pipeline-report Step 4 (SKILL.md, DESIGN-SYSTEM.md, signalforge.css)
- `prospect-research-multithreading` — invoked by comms-drafter, email-drafter, deal-strategy-coach
- `skill-orchestrator` — referenced by analysis-validator §11 and signalforge-feedback activation checklist
- Specialist validation skills in analysis-validator §12.4: `bonusly-data-questions`, `bonusly-product-questions`, `bonusly-business-reporting-questions`, `bonusly-rewards-questions`, `bonusly-ppp-questions`, `bonusly-feature-flag-questions`, `bonusly-deal-desk-questions`, `bonusly-datadog-questions`
- `caveman` — referenced in signalforge-claim-compressor ("Relationship to Caveman Skill")
Severity is CRITICAL because two of these (bonusly-brand, signalforge-reports) are blocking mandatory steps, not optional cross-references.
Proposal: REVIEW the deployment — either add these skills/files to the set and manifest, or annotate each reference as "external skill, resolved at runtime outside this bundle" so the manifest reflects reality. One proposal covering the whole dangling list.

====================================================
(4) VERSION CONFLICT
====================================================

FINDING 4.1 — WARNING — UPDATE_BODY
analysis-validator declares **Version: 3.6** in its header and changelog, but its own Section 7 validation-trail template hardcodes `Validator: analysis-validator v3.2`. Every published trail would stamp the wrong version. v3.6 should survive (header + changelog + pipeline-intelligence-report footer all say v3.6; v3.2 is a stale template artifact).
Proposal: UPDATE_BODY — change the trail template line to v3.6 (or make it a variable).

FINDING 4.2 — INFO — REVIEW
analysis-validator changelog lists v3.6 above v3.5, both dated May 9, 2026, while v2.6 (May 4) precedes v3.0–3.6 — ordering is descending-newest-first but 3.6/3.5 share a date with no sequence marker; minor ambiguity about which shipped last on May 9.
Proposal: REVIEW — add timestamps or sequence numbers to same-day changelog entries.

====================================================
(5) MANIFEST DESCRIPTIONS OVER 1,024 CHARACTERS
====================================================

Count: **0**.
Arithmetic: description_chars values are 656, 897, 996, 792, 965, 676, 945, 1004, 1006, 962, 1006, 708, 762, 656. Max = 1,006 (pipeline-intelligence-report and signalforge-claim-compressor, tied). 1,006 < 1,024, and no other value exceeds 1,006. Therefore zero rows exceed the limit.
FINDING 5.1 — INFO — no action required. (Note: pipeline-intelligence-report and signalforge-claim-compressor are within 18 characters of the cap — flag for awareness only.)

====================================================
(6) HARDCODED PAGE IDS, DATES, PERSON NAMES IN BODIES
====================================================

FINDING 6.1 — WARNING — UPDATE_BODY (person names)
- weekly-pipeline-report: title "Weekly Pipeline Report — Ben Lavin · Demand Generation · Bonusly"; "presented in chat for Ben's review"; "deliver the HTML file to Ben" — skill is person-bound.
- analysis-validator §12.3: full GTM roster with names + HubSpot owner IDs (Bryce Harmon, Hugo Lindqvist, Dana Mercer, Alex Franklin, Cole Ingram, Gavin Porter, Alaina Loori, Shealagh Coughlin, Ben Castelli, Amani Phipps, John Thomas, Yasmin Wahid, plus CSMs) — and pipeline-intelligence-report Phase 1 re-hardcodes AE owner IDs "verified May 2026," directly contradicting stale-pipeline-report Phase 2's rule "Never hardcode rep names or owner IDs. The AE roster changes."
- analysis-validator §G1-K/§10: escalation names "Manish or Amani."
- deal-strategy-coach ICP: routing names "Perseus" and "Farid."
- partner-digest: "Owner: Amani Phipps," Slack user ID `U03QLMBL7AR`, partner contact first names (Kelli, Jen Lee, Hani, Bryce, Sara).
- sales-forecast: "Alaina / VP Sales view" (changelog notes Elena → Alaina was already a hardcode fix — the pattern persists).
Proposal: UPDATE_BODY — replace person-bound rosters/owners with the dynamic owner-resolution pattern already proven in stale-pipeline-report Phase 2, and move named escalation contacts into a single maintained reference table.

FINDING 6.2 — WARNING — REVIEW (page/space/cloud/sheet/channel IDs)
- partner-digest: Cloud ID `73fe98de-a4a3-4869-9f8a-bb1eeed4cf7f`, Space ID `1958248479`, folder ID `2286616609`, canonical page IDs 2286321666, 2265382925, 2236940297, 2237825028, 2239365136, 2238283777.
- signalforge-feedback: page ID `2295136266`, parent `2234417154`, Build Log `2247295002`, spaceId 2232811524, same cloudId.
- sales-forecast: Space ID `2232811524`, Parent page ID `2232582148`, same cloudId.
- deal-strategy-coach: Confluence page `2257879045` (AE Excellence Playbook).
- weekly-pipeline-report: Google Sheets IDs `1CLZeOsElVDF_LF0ZG_t2nfwvhnZ6bpwqM_nX3WEYzcw` and `1ENuaEcCuLjdKhMvp8FK3Ys1ek5Aw9ZuOZhsHJJFoB_k`.
- stale-pipeline-report: Slack channel `#revops-team` ID `C0561C1JCPJ`; owner ID `55483190` (Bonusly Support).
- pipeline-intelligence-report / next-to-close: HubSpot portal ID `1973303` (used in every deal URL).
Proposal: REVIEW — consolidate IDs into one shared constants reference (they repeat across skills, e.g., the same cloudId in 3 skills) so a Confluence/portal migration doesn't require editing 6+ bodies.

FINDING 6.3 — WARNING — UPDATE_BODY (dates that will go stale)
- weekly-pipeline-report Step 0: "Business days complete in Q2 (April 1 – June 30, 2026)" and Step 2B "Q1 2026 context (static): Sales Bookings Actual: $365,152 vs. $475,000 plan (77%); Pipeline Addition Actual: $2,490,532 vs. $3,288,000 forecast (76%)" — quarter-locked in a skill that claims to run weekly year-round (sales-forecast v1.1 fixed exactly this problem for itself).
- model-selection registry: `last_checked: 2026-05-19`, deprecation date "April 14, 2026" — has a built-in 14-day staleness rule, so the date itself is a trigger, acceptable but will force self-update.
- deal-strategy-coach: pricing table labeled "2026"; Playbook title "April 2026."
- pipeline-intelligence-report: "v6 · May 2026," "AE owner IDs (verified May 2026)," "CONFIRMED STALE… last modified March 2023," "confirmed current as of May 4, 2026."
- closed-lost-analysis: dated examples throughout ("May 2026 sample," "MinIO: rep vacation May 4–12," "ai_closed_lost_reason field confirmed May 2026") plus named company examples (Softheon, Estee Lauder, LIFTOFF, Nestlé, Ozinga, Aurora Innovation, GCash, Ethos Cannabis, StickerYou).
- stale-pipeline-report: example dates "5/15," "5/19," "5/7"; changelog 2026-06-10.
- partner-digest: "May 16, 2026 issue as canonical reference," example week titles May 19 / June 2, 2026, "Pipeline Partner Plays Q2/Q3 2026."
- analysis-validator: Created April 26, 2026; Last Updated May 9, 2026; CALL_SPOTLIGHT_BRIEF removal "as of May 4, 2026"; DEALS table stale "as of March 28, 2023."
Proposal: UPDATE_BODY — highest priority on weekly-pipeline-report: make the quarter window computed (as sales-forecast v1.1 did) and move Q1-2026 static actuals to a reference file; for the rest, mark dated assertions with "verify at run time" where they function as live constraints (they mostly already are illustrative examples, which is acceptable).

====================================================
(7) MANIFEST DRIFT — BOTH DIRECTIONS
====================================================

FINDING 7.1 — INFO — no action
Files with no manifest row: 0. Files present: analysis-validator, closed-lost-analysis, comms-drafter, deal-strategy-coach, email-drafter, model-selection, next-to-close, partner-digest, pipeline-intelligence-report, sales-forecast, signalforge-claim-compressor, signalforge-feedback, stale-pipeline-report, weekly-pipeline-report = 14.
Manifest rows with no file: 0. Rows = same 14 names, and each declared_name matches its file column exactly.
Arithmetic: |files| = 14, |manifest rows| = 14, |files − manifest| = 0, |manifest − files| = 0.
Proposal: none required — manifest and file set are in sync. (Caveat: this reconciliation covers only the 14 provided files; the ~20 delegation targets in Finding 3.1 live outside both the file set and the manifest, so if the manifest is meant to cover the full runtime skill universe, it is incomplete by those entries.)

====================================================
SUMMARY
====================================================
CRITICAL: 3 (1.1 trigger dup comms/email, 1.2 trigger dup pipeline pair, 3.1 dangling targets)
WARNING:  6 (1.3 always-on collision, 2.1 delegation cycle, 4.1 v3.6/v3.2, 6.1 names, 6.2 IDs, 6.3 dates)
INFO:     3 (4.2 changelog order, 5.1 zero descriptions >1,024 chars, 7.1 no manifest drift)
runbook-from-slack-thread0.75
Plain-English question: Turn the resolved incident thread into a runbook.
Prompt — what the model saw
Turn the resolved incident thread into a runbook: ordered steps with the exact command or action taken, who ran it, how success was verified, and rollback for each step that changed state. Every step must trace to a message in the thread. Add nothing the thread does not contain; where the thread is unclear, mark the step as needing confirmation.
cost $0.0461 · 28s · in 2,971 / out 1,226 tok
✗ ordered_steps✓ rollback✓ needs_confirmation✓ sourced

✓ Correct Answer (ground truth)

Plain-English answer: Ordered steps from the thread, each with the exact command, who ran it, how success was verified, and rollback where state changed.
The answer in one line: rule = no steps beyond what the thread contains; every step traces to a message id.
ambiguous_steporder=4; msg=M05; mark=needs-confirmation
ruleno steps beyond what the thread contains; every step traces to a message id.
stepsorder=1; msg=M01; actor=Bryce Harmon; action=acknowledge PagerDuty alert, take IC; state_change=False; verified_by=M01 | order=2; msg=M02; actor=Farid Osman; action=check queue depth; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M02 (48,213 pending) | order=3; msg=M04; actor=Farid Osman; action=pause auto-enqueue via feature flag; command=bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'; state_change=True; rollback=bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'; verified_by=M07 (queue depth falling) | order=4; msg=M05; actor=Elena Sinclair; action=cleared the dead set; state_change=True; needs_confirmation=thread gives no command or method for clearing the dead set | order=5; msg=M06; actor=Bryce Harmon; action=scale reward-worker to 6 replicas; command=kubectl scale deployment/reward-worker --replicas=6; state_change=True; rollback=kubectl scale deployment/reward-worker --replicas=3; verified_by=M07/M08 | order=6; msg=M08; actor=Cole Ingram; action=verify queue drained and error rate baseline; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M08 (returns 0) … (+2 more)
graded checksordered_steps · rollback · needs_confirmation · sourced

Model's actual answer

show the model's full answer
RUNBOOK — incident-2026-08-29-reward-queue (resolved)
Source: incident_thread.md export, #incident-2026-08-29-reward-queue. Every step traces to a message ID. Nothing added.

STEP 1 — Acknowledge alert, take IC [M01]
Action: Bryce Harmon acknowledged the PagerDuty alert (reward-worker queue depth > 10k) and took IC.
Command: None documented.
Run by: Bryce Harmon.
Verified: No independent verification documented.
Rollback: None documented (no state change).

STEP 2 — Measure queue depth [M02]
Action: Assessed backlog.
Command: `bundle exec rake sidekiq:queue_depth`
Run by: Farid Osman.
Result reported: reward queue at 48,213 pending jobs; normal is under 500.
Rollback: N/A (read-only).

STEP 3 — Inspect dead set [M03]
Action: Farid Osman reported dead set has 112 jobs, all Redis::TimeoutError from around 13:58.
Command: Not documented in the thread — needs confirmation.
Run by: Farid Osman.
Rollback: N/A (read-only inspection as reported).

STEP 4 — Pause enqueue (STATE-CHANGING) [M04]
Action: Disabled auto-recognition enqueue to stop the bleed.
Command: `bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'`
Run by: Farid Osman.
Verified: No direct verification of the flag state is recorded in the thread — needs confirmation. (M07's falling queue depth is a later observation, not isolated proof of this step.)
Rollback (explicitly documented in M04): `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`

STEP 5 — Clear dead set (STATE-CHANGING, DESTRUCTIVE) [M05]
Action: Elena Sinclair cleared the dead set while in the console.
Command: Not documented — needs confirmation.
Run by: Elena Sinclair.
Verified: No verification documented — needs confirmation.
Rollback: Not documented — needs confirmation. (Dead-set clearing is state-changing and destructive; the thread does not establish a rollback, and this step should not be treated as an approved repeatable remediation merely because it happened during the incident.)

STEP 6 — Scale workers up (STATE-CHANGING) [M06]
Action: Scaled reward-worker deployment from 3 to 6 replicas.
Command: `kubectl scale deployment/reward-worker --replicas=6`
Run by: Bryce Harmon.
Verified: No direct replica-count verification recorded — needs confirmation. (M07: Farid Osman reported queue depth down to 9,400 and falling ~1,200/min; M08: Cole Ingram verified depth 0 — both are later queue observations, not isolated proof of this action's effect.)
Rollback (explicitly documented in M06): `kubectl scale deployment/reward-worker --replicas=3`

STEP 7 — Verify recovery [M08]
Action: Confirmed drain complete.
Command: `bundle exec rake sidekiq:queue_depth`
Run by: Cole Ingram.
Result: Returns 0. Error rate in Datadog back to baseline (reported by Cole Ingram).
Rollback: N/A (read-only verification).
Note: M07 (Farid Osman, 9,400 pending, falling ~1,200/min) is an interim report with no measurement command documented — needs confirmation.

STEP 8 — Re-enable enqueue (STATE-CHANGING) [M09]
Action: Re-enabled auto-recognition enqueue.
Command: `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`
Run by: Bryce Harmon.
Verified: Yes, directly — 40 new jobs processed cleanly in the next 3 minutes (per M09).
Rollback: Not documented — needs confirmation. (The thread does not record an approved rollback for this step; do not infer disabling enqueue as one.)

STEP 9 — Scale workers back down (STATE-CHANGING) [M10]
Action: Returned reward-worker deployment to 3 replicas.
Command: `kubectl scale deployment/reward-worker --replicas=3`
Run by: Bryce Harmon.
Verified: Queue stable at 0; incident declared resolved (per M10).
Rollback: Not documented — needs confirmation. (Do not infer scaling back to 6 as an approved rollback.)

UNRESOLVED GAPS (all need confirmation): dead-set inspection command (M03); exact command, verification, and rollback for dead-set clearing (M05); direct verification of flag state after disable (M04) and replica counts after each scale (M06, M10); measurement command behind M07; root cause of the Redis::TimeoutError burst — not addressed anywhere in the thread.
datadog-incident-timelinePASS 1.00
Plain-English question: Sequence the failure from the logs — what broke first, and what cascaded after it?
Prompt — what the model saw
From the log slice, sequence the failure: first error with timestamp, the cascade in order, the service and job involved, and the Datadog query that would confirm the first error. State what the logs do not show.
cost $0.0691 · 42s · in 3,701 / out 2,821 tok
✓ first_error_ts✓ service_named✓ cascade_ordered✓ query✓ gaps_stated

✓ Correct Answer (ground truth)

Plain-English answer: A reward-service Redis timeout at 14:01:12Z cascades to sidekiq at 14:01:40 and api-gateway at 14:03:05, with the Datadog query that confirms the first error.
The answer in one line: job_class_involved = RewardGiveJob
cascade_orderservice=reward-service; first_error_at=2026-09-03T14:01:12Z | service=sidekiq; first_error_at=2026-09-03T14:01:40Z | service=api-gateway; first_error_at=2026-09-03T14:03:05Z | service=web-app; first_error_at=2026-09-03T14:03:30Z
clean_servicespostgres
confirming_queryservice:reward-service status:error "@message:Redis::TimeoutError"
first_errortimestamp=2026-09-03T14:01:12Z; service=reward-service; message=Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s
job_class_involvedRewardGiveJob
logs_do_not_showroot cause of the Redis timeout, redis-server's own logs, infrastructure metrics (CPU/memory/network), any deploy or config change preceding the incident, who applied mitigation or when (no actor fields in logs)
graded checksfirst_error_ts · service_named · cascade_ordered · query · gaps_stated

Model's actual answer

show the model's full answer
FAILURE SEQUENCE (source: datadog_logs.csv + sidekiq_jobs.csv, 2026-09-03)

1. FIRST ERROR
2026-09-03T14:01:12Z — reward-service — ERROR — "Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s"
This is the earliest ERROR-level row in the slice; all rows before it (13:55:13Z–13:59:30Z) are INFO.

2. CASCADE, IN ORDER
- 14:01:12Z reward-service: Redis::TimeoutError to redis-primary:6379 (timeout after 5s) — origin
- 14:01:20Z / 14:01:30Z / 14:01:40Z reward-service: "retry exhausted for RewardGiveJob" (3 rows)
- 14:01:40Z sidekiq: "RewardGiveJob failed: Redis::TimeoutError; retrying in 60s"
- 14:01:46Z–14:01:57Z sidekiq_jobs.csv: J-00001–J-00006 (RewardGiveJob) failed_at, all Redis::TimeoutError
- 14:02:28Z sidekiq: RewardGiveJob failed; retrying
- 14:02:30Z sidekiq: WARN "Queue reward depth above 10,000"
- 14:02:36Z J-00013 (RecognitionDigestJob) fails — Redis timeout spreads to a second job class
- 14:02:51Z–14:02:58Z J-00007–J-00012 (RewardGiveJob) fail
- 14:03:05Z api-gateway: "502 upstream timeout calling reward-service /gives" — cascade reaches the edge
- 14:03:15Z / 14:04:55Z / 14:05:50Z J-00014–J-00016 (RecognitionDigestJob) fail
- 14:03:30Z web-app: "Give form submission failed: upstream 502 from api-gateway" — user-facing failure
- 14:03:31Z–14:06:47Z sidekiq: repeated RewardGiveJob failures (14:03:31, 14:04:22, 14:05:26, 14:06:47)
- 14:03:48Z / 14:04:13Z / 14:05:16Z / 14:06:52Z api-gateway: 502 upstream timeout calling reward-service
- 14:04:45Z / 14:05:42Z / 14:06:49Z web-app: Give form submission failed; upstream 502
- 14:22:10Z reward-service: "Redis connection restored; resuming job processing" — recovery
- 14:24:45Z sidekiq: "Queue reward depth below 500" — backlog drained

Arithmetic: 19 ERROR rows in datadog_logs.csv (reward-service 4, sidekiq 7, api-gateway 5, web-app 4 — 4+7+5+4=20 minus the 1 sidekiq WARN = 19 ERRORs + 1 WARN). Job failures: 16 total = 12 RewardGiveJob (J-00001–J-00012) + 4 RecognitionDigestJob (J-00013–J-00016). First error → Redis restored: 14:01:12Z to 14:22:10Z = 20m58s. First error → queue below 500: 14:01:12Z to 14:24:45Z = 23m33s.

3. SERVICE AND JOB INVOLVED
- Originating service: reward-service (dependent on redis-primary:6379)
- Worker: sidekiq
- Primary job: RewardGiveJob; secondarily RecognitionDigestJob
- Downstream: api-gateway (/gives) and web-app (Give form)

4. DATADOG QUERY TO CONFIRM THE FIRST ERROR
Log search, window bracketing 14:01:12Z:

service:reward-service status:error "Redis::TimeoutError" @timestamp:[2026-09-03T14:00:00Z TO 2026-09-03T14:02:00Z]

Expected confirmation: a hit at 2026-09-03T14:01:12Z with message "Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s", and no error-level hits for reward-service before that timestamp in the slice.

5. WHAT THE LOGS DO NOT SHOW
- Why Redis timed out: no logs from redis-primary itself (no redis service rows at all); root cause is absent.
- No remediation actions logged — nothing explains what caused "Redis connection restored" at 14:22:10Z (self-recovery vs. human/infra action is unknown).
- Final job outcomes: sidekiq_jobs.csv records failed_at only; whether J-00001–J-00016 eventually succeeded, dead-setted, or were dropped is not shown.
- Whether any rewards were double-given or lost (no idempotency/dedup evidence).
- Peak queue depth: only bounds are shown ("above 10,000" at 14:02:30Z, "below 500" at 14:24:45Z); the maximum and drain rate are not logged.
- Customer impact volume: number of failed Give form submissions beyond the 4 logged web-app ERROR rows is not shown.
- postgres rows (14:10:56Z–14:20:59Z, "checkpoint complete", INFO) show no errors; the logs neither implicate nor fully clear the database beyond that.
feature-flag-state-summaryPASS 1.00
Plain-English question: Summarize what each feature flag controls and who has it on.
Prompt — what the model saw
Summarize the state of the listed feature flags: what each controls per the code excerpt, which companies or segments have it on, targeting rules, and any flag with no code reference or no targets. Company counts must match the export. Do not describe a flag that is not in the export.
cost $0.0304 · 25s · in 773 / out 1,464 tok
✓ orphans✓ no_targets✓ counts_match✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: Recognition streaks on for 42 beta companies, redeem-flow redesign off with a 12-company target list, plus any flag with no code reference or no targets.
The answer in one line: rule = company counts must equal the export; do not describe flags absent from the export.
code_referencedrecognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, redeem_flow_redesign, analytics_dashboard_v3, ms_teams_app_v2
flags{"recognition_streaks_v2": {"state": "on", "targeting_rules": "segment:beta_companies", "company_count": 42}, "points_budget_guardrails": {"state": "on", "targeting_rules": "all_companies", "company_c
flags_with_no_targetspaused_offboard_cleanup
orphan_flags_no_code_referencelegacy_give_modal, survey_boosters_q3, paused_offboard_cleanup
rulecompany counts must equal the export; do not describe flags absent from the export.
graded checksorphans · no_targets · counts_match · no_invented

Model's actual answer

show the model's full answer
Feature flag summary (source: flags_export.csv + flag_code.md only; company counts taken verbatim from export)

FLAGS WITH CODE REFERENCES

1. recognition_streaks_v2 — ON
   - Controls: in app/models/recognition.rb, when enabled for a company, StreakTracker.record(give) runs — i.e., giving events are recorded into streak tracking.
   - Targeting: segment:beta_companies — 42 companies.

2. points_budget_guardrails — ON
   - Controls: in app/services/budget_service.rb, when enabled, BudgetService.new(company).enforce!(giver, points) runs — enforces point budget guardrails on each give.
   - Targeting: all_companies — 220 companies.

3. slack_dm_nudges — ON
   - Controls: in app/jobs/nudge_job.rb, the job returns early unless enabled; when enabled it calls SlackDm.send_nudge(user) — sends Slack DM nudges.
   - Targeting: segment:region_na — 87 companies.

4. redeem_flow_redesign — OFF
   - Controls: in app/controllers/redeem_controller.rb, enabled renders RedeemV2Component; else RedeemV1Component. Since the flag is OFF, the 12 targeted companies currently get the V1 redeem flow.
   - Targeting: targeted_list — 12 companies.

5. analytics_dashboard_v3 — ON
   - Controls: in app/controllers/analytics_controller.rb, when enabled sets @dashboard = AnalyticsV3.new(company) — serves the v3 analytics dashboard.
   - Targeting: segment:tier_three — 65 companies.

6. ms_teams_app_v2 — OFF
   - Controls: in app/services/teams_installer.rb, when enabled runs TeamsAppV2.install(company). Since OFF, the v2 Teams app is not installed for the 9 targeted companies.
   - Targeting: targeted_list — 9 companies.

FLAGS WITH NO CODE REFERENCE IN THE EXCERPT (behavior unknown from provided data)

7. legacy_give_modal — OFF, segment:legacy_plan, 14 companies. No usage found in flag_code.md; what it controls is not determinable from the given files.

8. survey_boosters_q3 — ON, segment:legacy_plan, 7 companies. No usage found in flag_code.md; what it controls is not determinable from the given files.

FLAG WITH NO TARGETS

9. paused_offboard_cleanup — OFF, targeting_rules is empty, company_count 0. Also no code reference. Effectively inert: off, untargeted, zero companies.

ARITHMETIC CHECKS
- ON flags (5): 42 + 220 + 87 + 65 + 7 = 421 company assignments.
- OFF flags (4): 12 + 9 + 14 + 0 = 35 company assignments.
- Total across all 9 flags: 421 + 35 = 456 company assignments (sums, not unique companies — segments may overlap; the export does not state uniqueness).
- Flags on segment:legacy_plan while OFF vs ON: legacy_give_modal off (14) and survey_boosters_q3 on (7) — both target the same segment with different counts; the export gives no explanation.

Not provided in the data: actual company names/aliases (export gives only segment names and counts), and code behavior for flags 7–9.
nda-triagePASS 1.00
Plain-English question: Triage each NDA green, yellow, or red — and quote the clause that decides it.
Prompt — what the model saw
Triage each NDA as GREEN (standard approval), YELLOW (counsel review), or RED (full legal review). Cite the clause that drives each classification by quoting it. Catch embedded non-solicits, non-competes, missing carve-outs, and unusual term or governing law. One line of reasoning per NDA.
cost $0.0748 · 22s · in 11,661 / out 722 tok
✓ green✓ yellow✓ red✓ non_solicit_caught✓ clause_quoted✓ correct_labels

✓ Correct Answer (ground truth)

Plain-English answer: NDA-1 green (standard mutual terms), NDA-2 yellow for an embedded 18-month non-solicit, with the driving clause quoted for every classification.
The answer in one line: non_solicit_in = NDA-2
checksclassification_per_nda, clause_quotes_exist_in_document, non_solicit_caught, carveouts_checked
classifications{"NDA-1": {"classification": "GREEN", "driving_reason": "standard mutual NDA: 2-year term, standard carve-outs, Delaware law", "driving_clause_quote": "Carve-outs: information that (a) is or becomes p
non_solicit_inNDA-2
graded checksgreen · yellow · red · non_solicit_caught · clause_quoted · correct_labels

Model's actual answer

show the model's full answer
NDA-1 — GREEN. Mutual, standard five carve-outs, Delaware law, 2-year term / 3-year survival, explicit no-license/no-exclusivity ("No license, no obligation to proceed, no exclusivity.") — nothing outside our standard NDA posture; no embedded restrictive covenants.

NDA-2 — YELLOW. Embedded non-solicit beyond confidentiality scope: "During the term of this Agreement and for eighteen (18) months thereafter, neither party shall solicit for employment or hire any employee of the other party with whom it came into contact under this Agreement." — 3-year total restriction (3-yr term + 18 months); mutual and contact-limited, so negotiable, but restrictive covenants are non-standard for an NDA and need counsel review before signature.

NDA-3 — RED. Multiple compounding issues: (1) a 3-year non-compete that would bar Bonusly from its own market — "For a period of three (3) years following the Effective Date, Recipient shall not, directly or indirectly, engage in or provide services to any business that competes with Discloser's business."; (2) foreign governing law + exclusive jurisdiction — "governed by and construed in accordance with the laws of the Republic of Ireland, and the parties submit to the exclusive jurisdiction of its courts."; (3) one-way obligations in favor of [PARTY C] with a 5-year term; (4) "No confidentiality carve-outs are stated in this Agreement" — missing all standard exclusions (public info, prior knowledge, independent development, legally compelled disclosure). Full legal review; recommend rejecting the non-compete outright and countering with our mutual template — never accept as drafted.

Note: parties are redacted, so no HubSpot/company context check was possible — classification is clause-based only.

Sound too good to be true? Ask Deal Desk to verify or consult with our external counsel.