← back to all models ← RevenueBench home
z-ai

z-ai/glm-5.3-flashx

Mean 0.978 · 36/40 perfect tests · $1.34 total · modalities in: text, image, video · out: text · each card: the prompt → the correct answer (gold) → this model's actual answer

What do the modalities mean?

Modalities in means what you can send this model: text, images, files, video, or audio. Modalities out means what it can send back. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what this model can handle in real work.

data-crm 5 tests
stage-hygiene-auditPASS 1.00
Plain-English question: How much of the open DS2+ pipeline is missing a substantive why-buys, and is anything missing DS1 timestamps or out of order?
Prompt — what the model saw
Audit stage hygiene on the open pipeline below. Rules: a deal at DS2 or later must have a substantive why-buys (15+ characters, not a placeholder such as tbd, n/a, see notes). A why_buys value of the form LONG(n) means the real text was present but truncated for display; treat it as present and substantive. Every deal should have a DS1 entry timestamp. A deal whose stage entry timestamps are out of order regressed (an earlier-stage timestamp is later than a later-stage timestamp). Report: percent of DS2+ deals non-compliant on why-buys overall and by owner (use the owners table for names), count of deals with no DS1 timestamp, count that regressed, and three process fixes. State definitions before numbers.
cost $0.0885 · 278s · in 60,343 / out 51,541 tok
✓ noncomp_pct✓ no_ds1_0✓ regressed_0✓ per_owner_breakdown✓ definitions_first✓ three_fixes

✓ Correct Answer (ground truth)

Plain-English answer: 81.5% of the 124 DS2+ deals are missing a substantive why-buys — Alex Franklin 48, Cole Ingram 19, Dana Mercer 18, Bryce Harmon 15, Farid Osman 1 — and every deal has a DS1 timestamp with none regressed.
The answer in one line: n_deals = 156; ds2plus = 124; noncomp_pct = 81.5; bryce_noncomp = 15; no_ds1 = 0; regressed = 0 (+1 more below)
n_deals156
bryce_noncomp15
ds2plus124
no_ds10
noncomp_pct81.5
regressed0
per_owner_noncompAlex Franklin=48 of 61 (78.7%); Cole Ingram=19 of 22 (86.4%); Dana Mercer=18 of 22 (81.8%); Bryce Harmon=15 of 18 (83.3%); Farid Osman=1 of 1 (100%)
graded checksnoncomp_pct · no_ds1_0 · regressed_0 · per_owner_breakdown · definitions_first · three_fixes

Model's actual answer

show the model's full answer
STAGE HYGIENE AUDIT — OPEN PIPELINE (156 deals)

DEFINITIONS
- Scope: all 156 open deals; DS2+ = deals at stage DS2, DS3, DS4, or DS5 (124 deals; 32 at DS1).
- Why-buys compliant (DS2+ only): why_buys is a non-placeholder string of 15+ chars. LONG(n) = present and substantive (compliant). Empty = non-compliant. No literal placeholder strings (tbd/n/a/see notes) exist anywhere in the file — every miss is a blank.
- No-DS1: t_ds1 field empty.
- Regressed: any earlier-stage timestamp strictly later than a later-stage one (same-day entries, e.g. t_ds1 = t_ds2, are not regressions per the stated rule).
- Data note: Deal-3EED2C has one extra empty field (19 vs 18 columns); under the only sensible alignment its why_buys is blank and timestamps are unaffected (counted non-compliant). Deal-57FF13 is missing source only.

NUMBERS
Why-buys non-compliance, DS2+ overall: 101 of 124 = 81.5% (101/124 x 100)
By stage: DS2 31/39 (79.5%), DS3 49/61 (80.3%), DS4 12/14 (85.7%), DS5 9/10 (90.0%)

By owner (DS2+ deals; owners with zero DS2+ deals excluded — Gavin Porter, Hugo Lindqvist, Elena Sinclair):
- Alex Franklin (84342457): 48/61 = 78.7%
- Dana Mercer (83155923): 18/22 = 81.8%
- Bryce Harmon (119337721): 15/18 = 83.3%
- Cole Ingram (83155924): 19/22 = 86.4%
- Farid Osman (716654662): 1/1 = 100.0% (n=1; not meaningful)
Within COMPLIANT deals, every pass is a LONG(n) value — no owner has typed a substantive why-buys manually.

COMMIT/BEST_CASE deals at DS2+ (forecast calls): 43/50 = 86.0% lack why-buys.

Deals with no DS1 timestamp: 0 of 156.
Regressed (out-of-order stage timestamps): 0.
Related observation (not a regression under the stated rule): 4 deals carry a t_ds4/t_ds3 stamp later than their current stage label (Deal-0D0211, Deal-BE3D9D, Deal-7E2131 labeled DS3 with t_ds4 set; Deal-C7F9BF labeled DS2 with t_ds3 set) — stage labels appear stale or deals were backed out without relabeling.

THREE PROCESS FIXES
1. Hard stage gate: block any transition into DS2+ unless why_buys passes the 15-char non-placeholder rule in CRM validation. Current misses are concentrated at DS2/DS3 entry (80 of 101), so the gate must fire at first promotion, not at forecast review.
2. Required field on forecast calls: why_buys mandatory for any deal placed in COMMIT or BEST_CASE (currently 86% blank) — no forecast submission in the weekly call without it.
3. Hygiene sweep + stage-label sync: weekly exception report listing DS2+ blanks (the 101 deals above) and deals whose highest stamp exceeds their stage label (the 4 above), worked to zero before pipeline review. Also backfill why-buys for COMMIT deals closing within 30 days first (e.g. Deal-547B2B, Deal-403845, Deal-B7EBD1 close 2026-09-10/11).
owner-resolutionPASS 1.00
Plain-English question: Every open deal has a blank owner name — resolve each one to a real person using the owners table.
Prompt — what the model saw
The deal extract below has an owner name field that is blank on nearly every row and an owner id field that is populated. Using the owners table, resolve each open deal to an owner name. Report: number of deals resolved, deals whose owner id has no match in the owners table (list them), the owner ids that map to archived/deactivated owners (list their names), and the total pipeline amount per resolved owner. Do not guess a name for any unmatched id.
cost $0.0425 · 95s · in 45,662 / out 18,104 tok
✓ resolved_count✓ unmatched_none✓ deactivated_flagged✓ jake_amount✓ justin_amount✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: All 156 deals resolve to six named owners; jake_amount maps to Bryce Harmon ($1,054,144) and justin_amount to Alex Franklin ($624,310).
The answer in one line: jake_amount = 1,054,144; justin_amount = 624,310; resolved = 156
jake_amount1,054,144
justin_amount624,310
resolved156
deactivated
per_ownerBryce Harmon=1,054,144; Alex Franklin=624,310; Dana Mercer=341,195; Cole Ingram=288,161.43; Farid Osman=4,134; Elena Sinclair=2,100
graded checksresolved_count · unmatched_none · deactivated_flagged · jake_amount · justin_amount · no_invented

Model's actual answer

show the model's full answer
Owner resolution complete — all 156 open deals parsed and joined against the 8-row owners table on owner_id.

1) DEALS RESOLVED
- 156 of 156 deals resolved to a named owner (every owner_id in deals_open.csv has an exact match in owners.csv).
- Owner-id → name mapping used: 119337721 = Bryce Harmon (35 deals), 84342457 = Alex Franklin (67), 83155923 = Dana Mercer (24), 83155924 = Cole Ingram (22), 716654662 = Farid Osman (7), 701163055 = Elena Sinclair (1). Deal-count check: 35+67+24+22+7+1 = 156.
- No duplicate deal_ids.

2) DEALS WITH NO MATCHING OWNER ID
- None. Every owner_id on the 156 rows matched an owners.csv row, so there is nothing to list and no name was guessed.

3) OWNERS THAT ARE ARCHIVED/DEACTIVATED AND HOLD OPEN DEALS
- None. No open deal maps to an archived owner.
- For completeness: owners.csv does contain two archived owners — Gavin Porter (1520255671) and Hugo Lindqvist (77260721) — but zero open deals carry either id.

4) TOTAL PIPELINE AMOUNT PER RESOLVED OWNER (sum of the amount column)

- Bryce Harmon (119337721), 35 deals = 1,054,144
  24000+19656+13500+7000+2520+240000+99000+72000+70000+63600+45000+1+21000+23400+13680+5502+8160+1+11400+1+36000+31500+6000+10800+30275+17400+12600+18000+37440+18828+2880+36000+20880+10920+25200 = 1,054,144

- Alex Franklin (84342457), 67 deals = 624,310
  14850+13770+11200+9000+6360+5400+3240+2484+1920+1080+7200+19000+2880+1400+4800+1632+10000+9300+2700+2160+1800+3600+3840+15000+1968+4000+3600+4800+3120+2520+9000+2400+62000+5400+5100+16700+4400+1620+2600+7200+18000+17000+8316+8100+18000+12600+24000+15000+9000+7200+3780+16200+7200+4680+1800+18000+2730+2400+3060+18000+12000+1800+4400+31200+7200+1600+60000 = 624,310

- Dana Mercer (83155923), 24 deals = 341,195
  11250+10500+9000+9000+5400+4800+4600+1920+15000+4200+18900+27000+43875+20000+60000+8100+16250+3150+5000+2100+23400+5400+7350+25000 = 341,195

- Cole Ingram (83155924), 22 deals = 288,161.43
  2760+1330+7225.40+3360+5616+2700+7781.20+6947.50+40000+12168+2480.40+11116+11193+1875+31750+32175+4779.88+3334.80+58529.25+4140+18000+18900 = 288,161.43

- Farid Osman (716654662), 7 deals = 4,134
  1249+1+1+1+2880+1+1 = 4,134

- Elena Sinclair (701163055), 1 deal = 2,100
  2100 = 2,100 (Deal-57FF13)

Grand total: 1,054,144 + 624,310 + 341,195 + 288,161.43 + 4,134 + 2,100 = 2,314,044.43 across all 156 open deals.

Data notes: amounts are used exactly as given (Cole Ingram's book includes four decimal amounts: 7225.4, 7781.2, 2480.4, 3334.8, 4779.88, 58529.25 — all included at face value). Six deals carry amount = 1 (placeholder-style); they are included as-is, not excluded or imputed.
gong-hubspot-join-integrityPASS 1.00
Plain-English question: How many call records point at deal ids that do not exist in the open pipeline?
Prompt — what the model saw
Check the integrity of the call-to-deal mapping below. The gong table lists calls in the last 90 days keyed to HubSpot deal ids; deals_open lists the current open default-pipeline deals. Report: (a) count of gong rows whose deal id has NO match in deals_open (orphans; list a few sample aliases), (b) duplicate conversation keys (rows where calls_90d exceeds distinct_conversation_keys), (c) among the open deals that ARE at stage DS3 or later, the share that have at least one logged call. Give sample aliases for the orphan finding, all taken from the data. Note: gong_owner_id is a Gong user id, not a HubSpot owner id.
cost $0.0715 · 176s · in 70,634 / out 34,856 tok
✓ orphan_count✓ dup_keys_0✓ ds3_coverage✓ sample_ids_present

✓ Correct Answer (ground truth)

Plain-English answer: 30 Gong call rows are orphans keyed to dead deal ids, none are duplicates, and DS3 call coverage is 29.4%.
The answer in one line: orphans = 30; ds3_cov_pct = 29.4; dup = 0
orphans30
ds3_cov_pct29.4
dup0
orphan_aliasesDeal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42, Deal-36EA09
graded checksorphan_count · dup_keys_0 · ds3_coverage · sample_ids_present

Model's actual answer

show the model's full answer
Ran the join programmatically (67 gong rows, 156 open deals, joined on hs_deal_id = deal_id; gong_owner_id not used as a join key).

**(a) Orphan gong rows — 30 of 67 (44.8%)**
Arithmetic: 67 gong rows, all with distinct deal ids; 30 have no match in deals_open → 30/67 = 44.8%.
Sample orphan aliases (all from data): Deal-8FA85D (60251290957, 46 calls), Deal-8FC3F9 (60251649055, 24), Deal-3B7945 (60251639682, 21), Deal-42B265 (61227242540, 21), Deal-9CCC42 (61430316324, 17), Deal-36EA09 (60251082126, 17).
Integrity signal: these 30 orphans include the highest call counts in the file, all carry calls_with_ds2_date = 0 / calls_with_sqo_date = 0, and their last_call dates run as late as 2026-11-25/2026-11-27 — outside the Sep-2026 horizon of every matched row (matched rows end by 2026-09-16). They look like a systematically distinct batch, not random loss.

**(b) Duplicate conversation keys — 0 rows**
Arithmetic: calls_90d > distinct_conversation_keys in 0 of 67 rows; all 67 rows have calls_90d exactly equal to distinct_conversation_keys. No duplicate conversation keys are detectable in this extract.

**(c) Call coverage among DS3+ open deals — 25 of 85 = 29.4%**
Arithmetic: deals_open at DS3/DS4/DS5 = 61 + 14 + 10 = 85. Of these, 25 have a match in the gong table (every gong row has calls_90d ≥ 3, so a match implies ≥1 logged call): DS3 12/61 (19.7%), DS4 6/14 (42.9%), DS5 7/10 (70.0%). Overall 25/85 = 29.4%.
Matched DS4/DS5 examples: Deal-D348E1, Deal-C26D20, Deal-547B2B, Deal-B7EBD1, Deal-A2B47C, Deal-C61CF7, Deal-584EE5, Deal-E53952, Deal-5408B0, Deal-944310, Deal-3974EB, Deal-C6D97A, Deal-25F752.
snowflake-sql-trapped-catalogPASS 1.00
Plain-English question: Write one Snowflake query for first-month usage per customer using only the provided, trap-laden catalog.
Prompt — what the model saw
Using ONLY the catalog provided, write one Snowflake SQL query returning, per customer company, for its first calendar month as a customer: unique givers, recognition count, successful redemption count. Read the catalog notes carefully: some tables are stale or unpopulated and must be avoided, and one documented filter (excluding deleted givers) must NOT be applied to historical giving counts. Put a comment above each FROM or JOIN explaining the table choice. Return only fenced SQL.
cost $0.0123 · 69s · in 981 / out 8,189 tok
✓ avoids_stale_deals✓ avoids_unpopulated_live✓ uses_cohort_summary✓ succeeded_filter✓ no_giver_deletion_filter✓ returns_sql

✓ Correct Answer (ground truth)

Plain-English answer: A single correct SQL query that uses the right catalog tables and avoids the trap columns — graded on the exact tables and columns.
The answer in one line: use_table = PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
avoid_stalePRODUCTION.HUBSPOT.DEALS (stale)
avoid_unpopulatedV2_LIVE.OBJECTS_DEALS (unpopulated)
forbidden_filterdeleted-giver exclusion must NOT be applied
outputone fenced SELECT with per-FROM/JOIN table-choice comment
required_filtersucceeded redemptions only
use_tablePRODUCTION.PLG.COMPANY_COHORT_SUMMARY
graded checksavoids_stale_deals · avoids_unpopulated_live · uses_cohort_summary · succeeded_filter · no_giver_deletion_filter · returns_sql

Model's actual answer

show the model's full answer
```sql
-- =============================================================================
-- Per customer company, first calendar month as a customer:
--   unique givers | recognition count | successful redemption count
-- =============================================================================
-- Catalog-driven choices:
--   * Customer anchor = calendar month of FIRST_SUB_PAYMENT_DATE in
--     PRODUCTION.PLG.COMPANY_COHORT_SUMMARY (the catalog's only per-company
--     customer-grain table). HubSpot deals were NOT usable as the anchor:
--     this catalog has no companies table / deal-to-company association, and
--     its deal tables include stale and unpopulated copies.
--     If "customer" should mean signup rather than first payment, swap
--     FIRST_SUB_PAYMENT_DATE for SIGNUP_DATE in customer_anchor.
--   * M1_GIVES / M1_REDEMPTIONS / M1_USERS deliberately NOT used: they are
--     cohort-relative aggregates (M1..M3 from signup), not calendar-month
--     figures, and M1_USERS is not a unique-giver count.
--   * Event source = PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2,
--     the catalog's documented source for redemption counts (schema name says
--     DEPRECATED, but the note says this is the source; finance-grade use
--     still requires confirmation per the note).
--   * Avoided per catalog: HUBSPOT_HUB_1973303.V2_LIVE.OBJECTS_DEALS
--     (unpopulated), PRODUCTION.HUBSPOT.DEALS (stale, last sync 2023-03),
--     PRODUCTION.HUBSPOT.GONG_HUBSPOT_MAP_FAST (retired).
--   * Documented rules applied: redemptions counted only when STATE =
--     'succeeded'; any deleted-giver exclusion is NOT applied (documented rule:
--     it understates historical giving) — no such filter appears below.
-- MISSING DATA (stated explicitly, not invented):
--   * The catalog documents only STATE in REDEMPTION_RECORDS_V2. Column names
--     company_id / giver_id / event_ts are placeholders that MUST be verified
--     against the live table before running.
--   * The company key in COMPANY_COHORT_SUMMARY is likewise unnamed in the
--     catalog (company_id placeholder below).
--   * No event-type column is documented, so recognition_count assumes one row
--     in REDEMPTION_RECORDS_V2 = one recognition event. If the table holds
--     redemption attempts only, recognition count cannot be produced from this
--     catalog (no other event-level giving table is documented).
-- =============================================================================

WITH customer_anchor AS (
    SELECT
        company_id,                                                    -- placeholder name; not documented in catalog
        DATE_TRUNC('month', FIRST_SUB_PAYMENT_DATE) AS first_customer_month
    -- Catalog: one row per self-serve company; the only per-company customer anchor provided.
    FROM PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
    WHERE FIRST_SUB_PAYMENT_DATE IS NOT NULL
),
monthly_events AS (
    SELECT
        e.company_id,                                                  -- placeholder name; not documented in catalog
        DATE_TRUNC('month', e.event_ts) AS event_month,                -- placeholder name; not documented in catalog
        COUNT(DISTINCT e.giver_id)      AS unique_givers,              -- placeholder name; not documented in catalog
        COUNT(*)                        AS recognition_count,          -- arithmetic: 1 row = 1 recognition event (assumption flagged above)
        COUNT_IF(e.state = 'succeeded') AS successful_redemption_count -- arithmetic: rows in month where STATE = 'succeeded' (documented rule)
    -- Catalog: documented source for redemption counts despite DEPRECATED schema name.
    -- No deleted-giver filter applied here on purpose (documented rule: it understates
    -- historical giving counts).
    FROM PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 e
    GROUP BY e.company_id, event_month
)
SELECT
    ca.company_id,
    COALESCE(me.unique_givers, 0)               AS unique_givers,
    COALESCE(me.recognition_count, 0)           AS recognition_count,
    COALESCE(me.successful_redemption_count, 0) AS successful_redemption_count
-- Keeps every customer company, including ones with zero month-1 activity (0/0/0).
FROM customer_anchor ca
-- Restricts each company's events to its first calendar month as a customer (month-equality join).
LEFT JOIN monthly_events me
    ON  me.company_id  = ca.company_id
    AND me.event_month = ca.first_customer_month
ORDER BY ca.company_id;
```
crm-completeness-enrichmentPASS 1.00
Plain-English question: Audit the CRM for missing and invalid fields, and propose the fixes.
Prompt — what the model saw
Audit the CRM extract for completeness and propose fixes. Deals need owner, stage, amount, close date, why-buys; companies need industry, employee count, HQ country; contacts need email, title, persona. Report completeness percent per field. Identify duplicate company clusters (name variants, shared domain) with a survivor per cluster. List invalid emails and domain mismatches. Fill missing company fields only where the enrichment export has a matching row; where CRM and enrichment disagree list both and recommend a source. Never invent a value. End with the 10 fixes with the most pipeline amount at stake.
cost $0.0411 · 188s · in 7,451 / out 27,694 tok
✓ dup_clusters✓ invalid_emails✓ survivor_language✓ completeness_pct✓ no_invent

✓ Correct Answer (ground truth)

Plain-English answer: 9 companies missing employee count, 6 missing HQ country, 4 invalid emails, 1 domain mismatch, 2 duplicate clusters, and 25 enrichment matches.
The answer in one line: n_contacts = 52; n_companies = 34; enrichment_matches = 25; invalid_email_count = 4; n_dup_clusters = 2; domain_mismatch_count = 1 (+4 more below)
n_contacts52
domain_mismatch_count1
enrichment_matches25
invalid_email_count4
n_companies34
n_dup_clusters2
company_missingindustry=0; employee_count=9; hq_country=6
domain_mismatchesCT-0011
duplicate_clusters['C-0A092931', 'C-0A092932'], ['C-0A092933', 'C-0A092934']
invalid_emailsCT-0010, CT-0080, CT-0081, CT-0192
graded checksdup_clusters · invalid_emails · survivor_language · completeness_pct · no_invent

Model's actual answer

show the model's full answer
# CRM Extract Audit

## 0. Deals: cannot be audited
No deals file was provided. Completeness for owner, stage, amount, close date, and why-buys = **not computable (0 deal rows supplied)**. No pipeline dollar amounts exist anywhere in the extract, so the final "top 10 fixes by pipeline at stake" is ranked by records affected as a proxy — no $ figures can be stated without inventing them.

## 1. Company completeness (34 rows)
| Field | Complete | % | Arithmetic |
|---|---|---|---|
| domain | 34/34 | 100.0% | 34/34 |
| industry | 34/34 | 100.0% | 34/34 — but only 13/34 (38.2%) use a clean canonical label (see §5) |
| employee_count | 25/34 | 73.5% | 25/34; 9 blank: C-EC3025, C-96039F, C-44EA29, C-D04904, C-B23205, C-60C75F, C-7BBDFA, C-50D386, C-93C8BF |
| hq_country | 28/34 | 82.4% | 28/34; 6 blank: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB |

## 2. Contact completeness (52 rows)
| Field | Complete | % | Arithmetic |
|---|---|---|---|
| contact_key / company_alias / domain | 52/52 | 100.0% | all FKs resolve to a company_alias |
| email | 48/52 usable | 92.3% | 52 populated, 4 malformed (§3): 48/52 |
| title | 39/52 | 75.0% | 39/52; 13 blank: CT-0000, CT-0022, CT-0072, CT-0080, CT-0081, CT-0092, CT-0120, CT-0121, CT-0122, CT-0132, CT-0141, CT-0162, CT-0170 |
| persona | 37/52 | 71.2% | 37/52; 15 blank: CT-0000, CT-0022, CT-0041, CT-0060, CT-0070, CT-0081, CT-0082, CT-0092, CT-0110, CT-0132, CT-0162, CT-0171, CT-0172, CT-0180, CT-0181 |

## 3. Invalid emails (4) and domain mismatch (1)
Invalid (truncated local@, no domain):
- CT-0010 / C-66D1FC — `user0@` (row domain 66d1fc.com)
- CT-0080 / C-92D97D — `user0@` (row domain 92d97d.com)
- CT-0081 / C-92D97D — `user1@` (row domain 92d97d.com)
- CT-0192 / C-425E2A — `user2@` (row domain 425e2a.com)

Domain mismatch:
- CT-0011 / C-66D1FC — `user1@other-domain.com` vs company/row domain 66d1fc.com

Likely fix for the 4 truncated: append the row's own domain (matches the `user0..2@<company domain>` pattern on sibling contacts) — but verify before writing. CT-0011: confirm whether other-domain.com is a legitimate secondary domain or a bad import.

## 4. Duplicate company clusters (by shared domain; aliases are opaque codes, so name-variant matching is not possible — no company names in the extract)
**Cluster 1 — acme-corp.com:** C-0A092931 (Technology, 500, US) and C-0A092932 (tech, 510, USA).
Survivor: **C-0A092931** (clean industry label; 500 employees). The 500 vs 510 employee conflict has no enrichment row to arbitrate (acme-corp.com absent from zoominfo_enrichment.csv) — flag for human review before deleting C-0A092932.
**Cluster 2 — globex.io:** C-0A092933 (SaaS, 200, US) and C-0A092934 (Technology, 200, US).
Survivor: **C-0A092934** (employee/country agree; fold "SaaS" in as an industry alias of the survivor). No enrichment row for globex.io either.
Neither cluster has any contacts attached, so merge risk is low.

## 5. Enrichment fills (only where zoominfo_enrichment.csv has a matching domain row; 25/32 distinct CRM domains covered; 7 absent: ba969b.com, 332637.com, 93c8bf.com, ee9ffb.com, c9bb20.com, acme-corp.com, globex.io)
Fillable employee_count (CRM blank, ZI populated) — 8 fills, lifting coverage 25/34 → 33/34 (97.1%):
C-EC3025←400, C-96039F←400, C-44EA29←400, C-D04904←400, C-B23205←400, C-60C75F←400, C-7BBDFA←400, C-50D386←400.
Not fillable: C-93C8BF (no ZI row).
hq_country blanks: **0 fillable** — all 6 blank-CRM companies also have blank/absent ZI country (C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5 blank in ZI; C-EE9FFB absent). Requires manual research; I am not filling them.

Disagreements (both values listed; recommend a source):
- Industry, 9 rows — CRM `Technology`/`tech`/`Tech ` vs ZI `Computer Software`: C-66D1FC, C-EC3025, C-44EA29, C-92D97D, C-D04904, C-77A95A, C-AA8DDA, C-B25F40, C-60C75F, C-425E2A. Recommend **ZoomInfo** as the taxonomy source (it is the enrichment system of record) — but note CRM values agree in meaning; this is a normalization problem.
- Country format, 10 rows — CRM `US`/`USA` vs ZI `United States` (C-66D1FC, C-950043, C-EC3025, C-96039F, C-77A95A, C-B23205, C-E51FB7, C-D0662E, C-425E2A, C-2D7423). Same value, different format; recommend full country names matching ZI.
- Employee counts where both present: 17/17 identical — no conflicts.
- CRM-internal conflict: acme 500 vs 510 (see §4) — no source available to arbitrate.

## 6. Top 10 fixes by records at stake (proxy — no deal amounts exist in the extract)
1. Backfill 15 missing personas (71.2% → 100%) — 15 contacts; persona drives segmentation/routing.
2. Backfill 13 missing titles (75.0% → 100%) — 13 contacts.
3. Standardize hq_country taxonomy — 18 non-conforming rows (`US`×9, `USA`×6, `UK`×3) across 28 populated companies; pick `United States`/`Canada`/`United Kingdom`.
4. Normalize industry taxonomy — 11 dirty rows (`tech`×4, `Tech `×4 incl. trailing space, `health care`×2, `SaaS`×1); align the 9 ZI-covered ones to `Computer Software`.
5. Fill 8 employee_counts from enrichment (73.5% → 97.1%) — 8 companies, 16 attached contacts.
6. Chase 14 contact-less companies (C-2C60E5, C-B97B4E, C-2D7423, C-7BBDFA, C-50D386, C-BA969B, C-332637, C-93C8BF, C-EE9FFB, C-C9BB20, C-0A092931, C-0A092932, C-0A092933, C-0A092934) — 14 of 34 companies (41.2%) unusable for outreach.
7. Repair 4 truncated emails (CT-0010, CT-0080, CT-0081, CT-0192) — blocks outreach at 3 companies (C-92D97D is 2/3 broken).
8. Merge acme-corp.com cluster onto C-0A092931 — resolves the 500 vs 510 employee conflict (needs human review; no enrichment arbiter).
9. Merge globex.io cluster onto C-0A092934 — resolves SaaS vs Technology.
10. Resolve CT-0011 domain mismatch (other-domain.com vs 66d1fc.com) — 1 contact; also pulls a non-existent domain out of C-66D1FC.

Residual gaps enrichment cannot close: 6 blank countries, C-93C8BF employee_count, 7 domains absent from the enrichment export, and all deal fields (no deals file).
deal-intelligence 5 tests
closed-lost-classificationPASS 1.00
Plain-English question: Why did we lose the last six months of deals — one primary category per deal, and whose side is the reason on?
Prompt — what the model saw
Classify every deal lost in the last 6 months below into one primary category: pricing, competitor, no decision, timing, product gap, champion left, or other. Also assign a side (Bonusly, buyer, unknown). Then summarize: category counts, the side split, the number of deals where the structured closed_lost_tag clearly disagrees with the free-text reason, and the two patterns most worth acting on. Base everything only on the tag and free-text provided.
cost $0.1023 · 490s · in 22,083 / out 71,848 tok
✓ total_90✓ competitor_top✓ category_counts_present✓ side_split✓ two_patterns

✓ Correct Answer (ground truth)

Plain-English answer: Every closed-lost deal gets exactly one category (pricing, competitor, no decision, timing, product gap, champion left, other) plus a side, per the computed classification.
The answer in one line: total_lost_6mo = 90; top_reason = competitor (26)
total_lost_6mo90
categoriespricing, competitor, no decision, timing, product gap, champion left, other
patternstwo most actionable, tag-vs-text disagreements counted
sideBonusly / buyer / unknown per deal
top_reasoncompetitor (26)
graded checkstotal_90 · competitor_top · category_counts_present · side_split · two_patterns

Model's actual answer

show the model's full answer
CLOSED-LOST CLASSIFICATION — 90 deals in file

DATA NOTES (missing/incomplete data, stated explicitly)
- The file's close dates span 2026-07-29 to 2026-09-30 (~2 months), not 6. If a full 6-month window was intended, earlier months were not provided.
- Only the structured tag and free-text reason were provided (per instruction). No conversation evidence, no AI-synthesized reason, and owner IDs have no names — so rep-level attribution and deeper root-cause work are not possible from this data.
- 21 rows have only "MIA/unresponsive/no response" as text — these are outcomes, not reasons. They are classified low-confidence below.

RULES APPLIED
- "no decision" = buyer never made a purchase decision: went dark (MIA), deprioritized/paused indefinitely, or failed internal approval.
- Timing = buyer wants it later and says so; Pricing = budget/price denial; Product gap = stated unmet capability; Champion left = decision-maker/sponsor loss; Competitor = chose/stayed with another vendor; Other = none of the above.
- Side: buyer = their decision/constraint/stated driver; Bonusly = a provided-field driver on our side; unknown = text gives no basis.

PER-DEAL CLASSIFICATION (alias exactly as given; $ = amount)

TIMING (20, $263,211)
Deal-DB0AAC $5,115 buyer — pause now, timeline to reconnect
Deal-91A056 $2,975 buyer — reconnect early 2027
Deal-29326C $6,300 buyer — "Timing"
Deal-831B7B $7,200 buyer — circle back in new year
Deal-39E25C $3,360 buyer — reconnect next year
Deal-B3ABED $40,001 buyer — revisit Q2 next year; budget for 2028
Deal-B6AC09 $3,000 buyer — revisiting 2027
Deal-E6E80A $24,000 buyer — pushed to early 2027
Deal-B038F0 $2,340 buyer — pushed to early 2027
Deal-175756 $2,880 buyer — on hold until 2027 (other priorities)
Deal-BB78F3 $6,600 buyer — plant survey items first; Bonusly still the end goal
Deal-15DA99 $19,600 buyer — bring back up early 2027
Deal-F4AF5D $5,760 buyer — early next year
Deal-79B7A1 $25,000 buyer — "Timing"
Deal-9F176A $54,600 buyer — pause until end of year
Deal-5E64CE $3,360 buyer — Nectar contract through Oct 2027, high exit fee; will move to Bonusly at contract end (tag said Doing-nothing/Cost; reclassified per text)
Deal-69CF3D $11,520 buyer — "On Hold"
Deal-ECBF89 $7,200 buyer — "On Hold for now"
Deal-D1A623 $25,200 buyer — "timing"
Deal-55867E $7,200 buyer — tag-led call; text is a soft decline ("at this time") with no re-engage date (weakest timing call in the set)

COMPETITOR (25, $382,234.96)
Deal-F7F635 $3,600 buyer — group went another direction (unnamed)
Deal-F97C37 $4,320 buyer — other vendor more diversified beyond R&R
Deal-422BA6 $3,000 buyer — competitor is preferred ADP TotalSource PEO partner (pre-built integrations)
Deal-381C8C $4,800 buyer — tag Competitor; text only says not moving forward, no vendor named
Deal-F1E8A6 $3,150 buyer — tag Competitor; text only says not moving forward, no vendor named
Deal-DDAB52 $4,000 buyer — Rippl: more at same cost, no FX friction
Deal-ACE061 $3,600 buyer — went another direction; rep believes HeyTaco (unconfirmed)
Deal-2D2F8D $4,800 buyer — different direction (unnamed)
Deal-0F96AA $76,800 buyer — RFP: not advanced to finalist stage (unnamed)
Deal-1BCA50 $15,000 buyer — stakeholder already far along with another vendor; budget/gift-card details
Deal-7CC678 $11,116 buyer — tag Competitor; text "Nothing specific provided."
Deal-242273 $60,000 buyer — other vendor better solved points-currency digitization / onsite spend
Deal-A2C349 $21,600 buyer — staying with Awardco, adding their surveys
Deal-C7156E $13,818 buyer — selected another vendor (unnamed)
Deal-8A0992 $7,336.56 buyer — chose a Canadian provider
Deal-D0C698 $2,000 buyer — returning to Kudos (past user)
Deal-EECC02 $66,690 buyer — went another direction (unnamed)
Deal-1E7DA9 $26,400 buyer — selected another platform (unnamed)
Deal-286F9C $13,860 buyer — went with another platform; "not really a good fit for us"
Deal-369281 $2,400 buyer — using what they have in Paylocity
Deal-9FCD0D $4,300 buyer — chose a Canadian company (CEO priority)
Deal-47F1A1 $10,004.40 buyer — staying with WorkTango another 12 months
Deal-BF2A98 $8,400 buyer — HiThrive recently deployed
Deal-64B19A $3,240 buyer — likely stayed with Motivosity
Deal-DC77FE $8,000 buyer — competitor offers more customization (label points as dollars); price explicitly not a factor

NO DECISION (31, $277,264.20)
MIA-tagged (21):
Deal-AC944F $3,400 unknown — "unresponsive"
Deal-214060 $2,880 unknown — "unresponsive"
Deal-21B045 $11,700 unknown — "MIA"
Deal-988493 $8,400 unknown — "mia"
Deal-F308CA $30,321 unknown — dark since April intro; ignored AE + ADR outreach
Deal-4664E1 $12,000 unknown — dark after intro; ignored AE + ADR outreach
Deal-D48E0B $14,931 unknown — "MIA"
Deal-583ADB $3,600 unknown — "MIA"
Deal-E0441F $2,405 Bonusly — stale deal inherited from a departed rep; no contact from consultant nor prospect (only row with an explicit Bonusly-side driver)
Deal-7CB44D $31,860 unknown — dark since demo; ignored AE + ADR outreach
Deal-AFA56C $3,000 unknown — unresponsive
Deal-D1AABF $23,400 unknown — no response
Deal-2BBA21 $2,310 unknown — dark since intro; ignored four nudges
Deal-386F6E $13,895 unknown — no response
Deal-3F86A0 $3,840 unknown — unresponsive
Deal-096750 $2,880 unknown — dark after intro; ignored four revivals
Deal-79E61A $7,020 unknown — unresponsive
Deal-AE7C4E $2,800 unknown — unresponsive
Deal-DAB4F1 $3,450 unknown — unresponsive
Deal-B4B50F $21,060 unknown — unresponsive
Deal-5885B9 $7,200 unknown — "MIA"
Doing-nothing-tagged (8):
Deal-13E9CF $33,750 buyer — R&R deprioritized by org (text explicitly says NOT budget); reach out next year
Deal-E74A73 $2,100 buyer — will test points manually first; maybe next year
Deal-50E5D8 $4,800 buyer — leadership pause; will reach out if that changes
Deal-413C56 $2,760 buyer — back-to-school priority; CEO not ready
Deal-FEDBCB $2,000 buyer — reconnect end of year; "not super engaged"
Deal-7FBAC6 $7,200 buyer — leadership paused "(again)"
Deal-2FEDDB $2,200 buyer — unsure she can get it moving
Deal-ABD14C $5,002.20 buyer — not interested in the program
Lost-DM-tagged (2):
Deal-70F704 $3,000 unknown — narrowed to anniversary-award automation only, then MIA (tag=Lost DM doesn't match text)
Deal-FAC17C $2,100 buyer — contract out 2 months; no final approval from Executive IT Director

PRICING (5, $172,450)
Deal-7ED004 $60,000 buyer — no budget approval
Deal-C33D91 $7,200 buyer — significant budget cuts; not approved
Deal-DAFB82 $30,000 buyer — budget needed for other priorities; not budgeted until 2028; will loop back
Deal-8A119B $3,250 buyer — didn't get approval
Deal-7B2236 $72,000 buyer — budget + wants "simpler and cheaper" (tag Doing-nothing/Cost; pricing per text)

PRODUCT GAP (4, $118,245)
Deal-9048EB $41,790 buyer — "bad fit based on their desired setup and multiple feature gaps" (tag said MIA)
Deal-3618CC $15,600 buyer — "Wanted Surveys" (tag said Lost DM)
Deal-5AD03E $24,000 buyer — "Wanted more defined budget access" (tag said Competitor)
Deal-981AD4 $36,855 buyer — UI fit + "not UK focused" (tag Feature Request, consistent)

CHAMPION LEFT (2, $16,740)
Deal-ED9AE7 $2,340 unknown — tag Lost DM; text "Timing, budget, authority" with no ranking; side unknowable from text
Deal-F325A5 $14,400 buyer — layoffs + change in leadership; no longer a priority

OTHER (3, $37,800)
Deal-5DB9B0 $10,800 unknown — "Spam." (record was junk, not a real loss)
Deal-2A292B $6,000 buyer — building something simple internally
Deal-8E27DA $21,000 buyer — moved forward with a swag provider only; didn't want R&R currently

SUMMARY

Category counts (n/90, % = n/90; ARR share of $1,267,945.16 total):
- No decision: 31 (34.4%) — $277,264.20 (21.9%)
- Competitor: 25 (27.8%) — $382,234.96 (30.1%)
- Timing: 20 (22.2%) — $263,211.00 (20.8%)
- Pricing: 5 (5.6%) — $172,450.00 (13.6%)
- Product gap: 4 (4.4%) — $118,245.00 (9.3%)
- Other: 3 (3.3%) — $37,800.00 (3.0%)
- Champion left: 2 (2.2%) — $16,740.00 (1.3%)
Check: 31+25+20+5+4+3+2 = 90.

Side split:
- Buyer: 67 (67/90 = 74.4%) = competitor 25 + timing 20 + no decision 10 + pricing 5 + product gap 4 + champion-left 1 (Deal-F325A5) + other 2
- Unknown: 22 (22/90 = 24.4%) = 20 pure-MIA rows + Deal-5DB9B0 + Deal-ED9AE7
- Bonusly: 1 (1/90 = 1.1%) = Deal-E0441F only
Caveat: the 1.1% Bonusly-side share reflects what these two thin fields can show, not a clean bill of health — 21 deals died unattended, which usually hides rep-side engagement failures this data can't confirm.

Tag vs free-text disagreements: 5 clear
1. Deal-5DB9B0 — "Does not fit ICP" vs "Spam."
2. Deal-70F704 — "Lost DM" vs narrowed use case + MIA (no DM loss in text)
3. Deal-3618CC — "Lost DM" vs "Wanted Surveys"
4. Deal-5AD03E — "Competitor" vs "Wanted more defined budget access" (feature need, no competitor)
5. Deal-8E27DA — "Feature Request" vs bought swag only / didn't want R&R (no feature request in text)
Borderline (not counted): Deal-13E9CF (tag includes Cost; text says explicitly not budget), Deal-9048EB (tag MIA; text names feature gaps), Deal-381C8C / Deal-F1E8A6 / Deal-7CC678 (Competitor tag, text silent), Deal-2A292B (doing-nothing vs internal build), Deal-55867E (timing vs undated decline).

Two patterns most worth acting on
1. Silent deaths. 21 deals (23.3%) worth $212,352 (16.8%) closed as MIA/unresponsive with no stated cause; several texts explicitly record "no meaningful contact since demo/intro — ignored outreach from me and the ADR." These produce zero loss intelligence and signal the deal died unattended. Action: a no-meaningful-contact escalation SLA (e.g., 48h post-demo follow-up, ~2-week dark-deal flag before closed-lost), require a stated reason before closing as lost, and multi-thread before demo. Deal-E0441F shows the failure mode when a rep departs with no handoff.
2. Dated timing losses with no system to catch them. 20 deals (22.2%) worth $263,211 (20.8%) almost all name a specific window — "early 2027," "Q2 next year," "end of year," Oct 2027 contract end — and two put the burden on us ("will reopen if they reach back out" / "if things change"). Uncycled timing is how deals become Pattern 1. Action: every timing loss leaves with a dated, calendared re-engage task and an owner; "on hold" without a date is treated as lost.
Combined, patterns 1+2 = 41 deals (45.6%) and $475,563 (37.5%) of lost ARR. Watch item: competitor is the largest ARR bucket ($382,235, 30.1%) and stated differentiators cluster on breadth/integrations/customization (diversified offerings, ADP PEO integrations, points-as-dollars labeling, onsite points spend) — price was explicitly not the factor where stated. Competitor named in only 8 of 25 competitive losses (2 hedged, 17 unnamed) — tag hygiene would sharpen this.
pipeline-tieringPASS 1.00
Plain-English question: Score all 156 open deals into exactly one of LOCK / ACTION / BUILD / REVIVE / WATCH / RISKY.
Prompt — what the model saw
Score the open pipeline below on these signals: forecast category, stage, meetings_30d, contact count, and engagement recency. Assign exactly one tier per deal: LOCK, ACTION, BUILD, REVIVE, WATCH, or RISKY (RISKY = the forecast category disagrees with the engagement evidence). Never tier a deal LOCK if it has zero meetings_30d. Do NOT list every deal. Return ONLY a compact JSON object with exactly these keys: - "tier_counts": object mapping each tier to the number of deals in it (must sum to the total deal count), - "tier_examples": object mapping each tier to up to 3 example deal aliases from the data, - "risky_deals": array of the deal aliases you tiered RISKY, - "lock_violations": number of deals tiered LOCK that have zero meetings_30d (should be 0), - "pipeline_shape": one short paragraph on the overall pipeline shape. Note: inbound_emails_30d is 0 for every row due to a data defect, so treat meetings_30d as the inbound signal.
cost $0.0804 · 295s · in 38,444 / out 48,832 tok
✓ json_returned✓ tier_vocabulary✓ tier_counts_sum✓ risky_deals_listed✓ no_lock_violation_claimed

✓ Correct Answer (ground truth)

Plain-English answer: All 156 open deals tiered exactly once across the six tiers using forecast category, stage, meetings, contacts, and recency.
The answer in one line: total = 156
total156
graded checksjson_returned · tier_vocabulary · tier_counts_sum · risky_deals_listed · no_lock_violation_claimed

Model's actual answer

show the model's full answer
```json
{
  "tier_counts": {
    "LOCK": 5,
    "ACTION": 35,
    "BUILD": 68,
    "REVIVE": 13,
    "WATCH": 28,
    "RISKY": 7
  },
  "tier_examples": {
    "LOCK": ["Deal-D348E1", "Deal-C26D20", "Deal-403845"],
    "ACTION": ["Deal-25F752", "Deal-E53952", "Deal-4F775F"],
    "BUILD": ["Deal-C6FE92", "Deal-D73B89", "Deal-B25F40"],
    "REVIVE": ["Deal-2D1F1B", "Deal-66D1FC", "Deal-950043"],
    "WATCH": ["Deal-EC3025", "Deal-92D97D", "Deal-44EA29"],
    "RISKY": ["Deal-7BBDFA", "Deal-547B2B", "Deal-B7EBD1"]
  },
  "risky_deals": [
    "Deal-2465CE",
    "Deal-547B2B",
    "Deal-584EE5",
    "Deal-A2B47C",
    "Deal-B7EBD1",
    "Deal-FD9F4E",
    "Deal-7BBDFA"
  ],
  "lock_violations": 0,
  "pipeline_shape": "Of 156 open deals ($2,314,044), only $40,039 (5 deals) is LOCK — all COMMIT with ≥1 meeting_30d and contact inside 7 days — while 63% of deals (99/156) show zero meetings_30d. Six of seven DS5 COMMIT deals with no meetings_30d land in RISKY (Deal-2465CE, Deal-547B2B, Deal-584EE5, Deal-A2B47C, Deal-B7EBD1, Deal-FD9F4E), plus Deal-7BBDFA (BEST_CASE, silent 45+ days); total overstated COMMIT/BEST_CASE exposure ≈ $75,330. The book is top-heavy in early stage: 68 BUILD (DS1–DS3 with fresh meetings, $797,947) and 28 WATCH ($618,694, mostly owner 119337721 PIPELINE DS1 deals incl. 8 $1-placeholder amounts), while $529,040 sits in REVIVE — large DS1/DS2 PIPELINE deals (Deal-2D1F1B $240k, Deal-66D1FC $99k, Deal-950043 $70k) untouched 15+ days — so near-term coverage depends on reactivating REVIVE/WATCH rather than the thin LOCK layer. Caveats: Deal-3EED2C and Deal-57FF13 have no engagement rows (data missing) and were tiered on stage/close-date signals only; Deal-333EBB (close 2026-08-28) and Deal-57FF13 (2026-09-02) are past close date; inbound_emails_30d = 0 everywhere per the stated defect, so meetings_30d was the sole inbound signal."
}
```
call-transcript-extractionPASS 1.00
Plain-English question: Extract the CRM write-back fields from each call transcript as JSON.
Prompt — what the model saw
For each transcript, extract CRM write-back fields as JSON: why-buys (prospect statements only), pain points, stakeholders from the speaker list, budget signal (prospect-stated or null), timeline signal, competitor mentioned (only if the prospect raised it), next step (only if explicitly agreed), objections, confidence. Never fill a field from what the rep said or hoped.
cost $0.0122 · 57s · in 1,897 / out 7,812 tok
✓ json_emitted✓ budget_signal✓ null_discipline✓ stakeholders✓ competitor_prospect_only

✓ Correct Answer (ground truth)

Plain-English answer: Per transcript: prospect-only why-buys, pain points, stakeholders, budget and timeline signals, and competitors — e.g. TX-001: '$40k earmarked', live before open enrollment.
The answer in one line: rep = Alex Franklin
allowed_newTX-001, TX-002, TX-003, TX-004, TX-005, TX-006, Achievers, Bucketlist, Nectar
per_transcript{"TX-001": {"why_buys": ["automating anniversary and birthday awards"], "pain_points": ["HR team of three cannot keep up with it manually", "track everything in a spreadsheet, and people slip through
repAlex Franklin
rep_sourced_nulls{"TX-003": ["budget_signal \u2014 only the rep mentioned pricing/budget"], "TX-004": ["next_step \u2014 rep proposed, prospect said 'no promises'"], "TX-006": ["budget_signal \u2014 rep decoy, prospec
graded checksjson_emitted · budget_signal · null_discipline · stakeholders · competitor_prospect_only

Model's actual answer

show the model's full answer
```json
[
  {
    "transcript_id": "TX-001",
    "deal_alias": "Deal-CFE7F4",
    "why_buys": [
      "The big win for us would be automating anniversary and birthday awards — our HR team of three cannot keep up with it manually. (Prospect, VP People)",
      "Right now we track everything in a spreadsheet, and people slip through the cracks. (Prospect, HR Admin)"
    ],
    "pain_points": [
      "Anniversary/birthday awards handled manually; HR team of three cannot keep up",
      "Recognition tracked in a spreadsheet; people slip through the cracks"
    ],
    "stakeholders": ["Prospect (VP People)", "Prospect (HR Admin)"],
    "budget_signal": "About $40k earmarked for engagement tools this fiscal year (stated by VP People)",
    "timeline_signal": "Wants it live before open enrollment in November",
    "competitor_mentioned": "Achievers — raised by prospect; evaluated last year, rejected as 'too heavy for a team our size' (historical, not active)",
    "next_step": "Security review with prospect's IT lead on September 12 (explicitly agreed by VP People)",
    "objections": [
      "SSO and audit logs required for IT to sign off (HR Admin)"
    ],
    "confidence": "High — prospect-stated budget, target date, and an explicitly agreed, dated next step"
  },
  {
    "transcript_id": "TX-002",
    "deal_alias": "Deal-70BB30",
    "why_buys": [
      "We want to tie recognition to retention for our hourly workforce — regretted turnover there is over 30%. (Prospect, Head of Total Rewards)",
      "Integration with Workday has to be rock solid — that's my one condition. (Prospect, CFO)"
    ],
    "pain_points": [
      "Regretted turnover in the hourly workforce is over 30%; recognition not currently tied to retention"
    ],
    "stakeholders": ["Prospect (Head of Total Rewards)", "Prospect (CFO)"],
    "budget_signal": "$25k pilot budget approved by finance for this quarter (stated by CFO)",
    "timeline_signal": "Decision wanted by end of September",
    "competitor_mentioned": null,
    "next_step": "Rep sends pilot agreement; CFO routes it to legal this week (explicitly agreed)",
    "objections": [
      "Workday integration must be 'rock solid' — stated as the CFO's one condition (gating requirement)"
    ],
    "confidence": "High — approved pilot budget, decision deadline, agreed next step; prospect stated this is the first real vendor demo"
  },
  {
    "transcript_id": "TX-003",
    "deal_alias": "Deal-530B50",
    "why_buys": [
      "We need to make recognition visible across our 12 retail locations. (Prospect, People Ops Manager)"
    ],
    "pain_points": [
      "Recognition not visible across 12 retail locations",
      "Store managers have zero budget autonomy for on-the-spot recognition today"
    ],
    "stakeholders": ["Prospect (People Ops Manager)"],
    "budget_signal": null,
    "timeline_signal": "No rush on prospect side until Q1",
    "competitor_mentioned": "Bucketlist — raised by prospect: 'My CEO used Bucketlist at her last company and liked it' (positive affinity, historical)",
    "next_step": "Call with the CEO — prospect will send two time slots (explicitly agreed, undated)",
    "objections": [
      "CEO must be sold first and decides anything people-related (economic buyer not yet engaged)",
      "No urgency until Q1"
    ],
    "confidence": "Low — no stated budget, decision-maker unengaged, explicit no-rush; next step agreed but undated"
  },
  {
    "transcript_id": "TX-004",
    "deal_alias": "Deal-180D02",
    "why_buys": [
      "We want to consolidate three separate recognition tools into one. (Prospect, VP People)",
      "We're paying for three tools and none of them talk to our HRIS. (Prospect, VP People)"
    ],
    "pain_points": [
      "Three separate recognition tools in use, none integrated with the HRIS",
      "Paying for three redundant tools",
      "Procurement/security friction: cycle runs six to eight weeks minimum; last vendor's security review took three months (IT Security Lead)"
    ],
    "stakeholders": ["Prospect (VP People)", "Prospect (IT Security Lead)"],
    "budget_signal": "Approval threshold: under $15k annually can be approved by VP People without going to the board (no commitment made)",
    "timeline_signal": "Procurement cycle runs six to eight weeks minimum; no target close date stated",
    "competitor_mentioned": null,
    "next_step": null,
    "objections": [
      "Security-review history: last vendor's review took three months — IT Security Lead's stated hesitation",
      "CFO follow-up not agreed: 'Maybe — I need to check her calendar, no promises'"
    ],
    "confidence": "Low — no agreed next step, six-to-eight-week procurement floor, security-review hesitation"
  },
  {
    "transcript_id": "TX-005",
    "deal_alias": "Deal-F8767A",
    "why_buys": [
      "Two things: automate service milestones, and give us analytics on recognition equity across departments. (Prospect, HR Director)",
      "Our night-shift teams feel invisible — their engagement scores run 20 points lower. (Prospect, People Ops Coordinator)"
    ],
    "pain_points": [
      "Night-shift teams feel invisible; engagement scores run 20 points lower",
      "Service milestones not automated",
      "No analytics on recognition equity across departments",
      "Exec team skeptical after a failed rollout two years ago"
    ],
    "stakeholders": ["Prospect (HR Director)", "Prospect (People Ops Coordinator)"],
    "budget_signal": "$12k approved under the engagement line (stated by HR Director)",
    "timeline_signal": "Must be running before the January all-hands",
    "competitor_mentioned": "Nectar — raised by prospect; actively mid-pilot, vendor 'would need to beat that experience'",
    "next_step": "Rep to present to the exec team on October 2 (explicitly agreed by HR Director)",
    "objections": [
      "Active Nectar pilot — incumbent experience to beat",
      "Exec team skepticism after a failed rollout two years ago"
    ],
    "confidence": "Medium — approved budget, dated deadline, and dated exec presentation agreed; active competitor pilot and exec skepticism are the risks"
  },
  {
    "transcript_id": "TX-006",
    "deal_alias": "Deal-EE195F",
    "why_buys": [
      "I'd love to cut the admin time on service awards. (Prospect, HR Manager)",
      "I personally spend five hours a month ordering and shipping plaques. (Prospect, HR Manager)"
    ],
    "pain_points": [
      "Five hours per month of admin time on ordering and shipping plaques",
      "Service-awards process is admin-heavy"
    ],
    "stakeholders": ["Prospect (HR Manager)"],
    "budget_signal": "No figure stated; 'Budget isn't the issue — time is' (HR Manager)",
    "timeline_signal": "Fiscal year resets in February; prospect says a Q1 start is realistic",
    "competitor_mentioned": null,
    "next_step": "Rep sends one-page overview; HR Manager forwards it to the COO this week (explicitly agreed)",
    "objections": [
      "Build-vs-buy risk: 'our COO usually prefers building things in-house' (COO not yet engaged)"
    ],
    "confidence": "Medium — quantified pain and agreed next step from an engaged contact, but no budget figure, soft Q1 timeline, and an unengaged COO with an in-house build preference"
  }
]
```
next-to-closePASS 1.00
Plain-English question: Which 3 deals are closest to signature right now?
Prompt — what the model saw
Which 3 deals are closest to signature right now? Use the CRM data and the Slack thread; Slack often has the real status before the CRM does. For each: deal, amount, why it is close, what is left. Under 250 words.
cost $0.0036 · 16s · in 1,232 / out 1,108 tok
✓ exactly_three✓ top3_named✓ slack_first✓ blocked_excluded

✓ Correct Answer (ground truth)

Plain-English answer: Deal-547B2B ($11,200), Deal-B7EBD1 ($9,000), and Deal-403845 ($9,000) — all DS5 COMMIT deals closing September 10-11, with what is left on each.
blocked_excludeddeal_alias=Deal-2465CE; reason=champion left + procurement freeze (Slack #sales-sparring 2026-09-04)
exactly_three1
ranking_ruleSlack signature-imminent deal ranks first; remaining ranked by stage (DS5>DS4), forecast (COMMIT>BEST_CASE>PIPELINE), close_date asc, amount desc
top3rank=1; deal_alias=Deal-547B2B; amount=11200; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11 | rank=2; deal_alias=Deal-B7EBD1; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-10 | rank=3; deal_alias=Deal-403845; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11
graded checksexactly_three · top3_named · slack_first · blocked_excluded

Model's actual answer

show the model's full answer
**Top 3 closest to signature** (CRM DS5/COMMIT + corroborating Slack signal, 2026-09-04):

1. **Deal-547B2B** — $11,200, Alex Franklin, close 2026-09-11.
   Why close: Slack — "redlines came back clean this morning, the signing page is out, and their VP People told me they are signing tomorrow. This one is signature-imminent."
   What's left: counterparty signature only.

2. **Deal-403845** — $9,000, Alex Franklin, close 2026-09-11.
   Why close: Slack — "moving fine... the order form is with their finance team."
   What's left: finance approval of the order form, then signature.

3. **Deal-A2B47C** — $6,360, Alex Franklin, close 2026-09-11.
   Why close: Slack — "still warm, just normal legal-review pace." DS5/COMMIT in CRM.
   What's left: legal review clearance, then signature.

Combined value of the 3: 11,200 + 9,000 + 6,360 = **$26,560**.

**Exclusions and caveats:**
- **Deal-2465CE** ($5,400, DS5/COMMIT, 2026-09-10) is excluded despite CRM status: Slack says the champion left, procurement froze new vendors, Dana is "pulling it out of commit... realistically a Q4 deal now."
- **Deal-B7EBD1** ($9,000, DS5/COMMIT, 2026-09-10) is the nearest-dated commit but has no Slack signal to verify it; it would be 4th on evidence available.
- The three selected all carry the same 2026-09-11 close date; ranking is by Slack proximity-to-signature (signature-imminent > order form at finance > legal in progress).
product-gap-detectionPASS 1.00
Plain-English question: Which prospect complaints are true product gaps versus rollout or enablement issues?
Prompt — what the model saw
Surface true product gaps raised by prospects in active deals. For each candidate: quote the prospect line, classify as real gap, plan gate, or rollout/enablement issue using the product docs excerpt, and name the deal and amount. Only real gaps go in the summary. Rep language does not count as prospect voice.
cost $0.0044 · 22s · in 780 / out 1,928 tok
✓ real_gaps✓ plan_gate_classified✓ enablement_classified✓ gap_summary

✓ Correct Answer (ground truth)

Plain-English answer: Two real gaps (SCIM provisioning, ADP Workforce Now integration) separated from plan gates and enablement issues, each with the prospect quote and deal.
allowed_newTG-001, TG-002, TG-003, TG-004, TG-005, SCIM user provisioning, ADP Workforce Now integration, Custom report builder
enablement_issuestranscript_id=TG-004; deal_alias=Deal-84DBA6; amount=16000; feature=Slack integration recognition adoption; quote=We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.; why_enablement=Slack integration exists on all plans; managers were never trained
excluded_rep_voicetranscript_id=TG-005; line=Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap.; reason=rep voice, not prospect-raised
plan_gatestranscript_id=TG-003; deal_alias=Deal-CFE7F4; amount=32000; feature=Custom report builder; tier_required=Enterprise; quote=I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?
real_gapstranscript_id=TG-001; deal_alias=Deal-EC3025; amount=62000; feature=SCIM user provisioning; quote=We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.; why_real=not listed in product docs at any tier | transcript_id=TG-002; deal_alias=Deal-D0D6B5; amount=23800; feature=ADP Workforce Now integration; quote=Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.; why_real=supported HRIS list is Workday/BambooHR/Gusto/Rippling only
graded checksreal_gaps · plan_gate_classified · enablement_classified · gap_summary

Model's actual answer

show the model's full answer
**Product-gap screen — transcripts_gaps.csv vs product_docs.md**

Amount note: no deal amounts exist anywhere in the provided files (no amount column), so amounts are unavailable for all deals below. Stated per instruction, not omitted silently.

Screened 5 transcript candidates; 4 carry prospect voice (rep lines excluded per instructions).

**Candidate-by-candidate:**

1. REAL GAP — Deal-EC3025
   - Prospect (IT Security Lead): "We need SCIM user provisioning — our security team will not approve any tool that requires manual account management."
   - Docs: "SCIM user provisioning ... NOT currently listed as supported capabilities" (absent from every tier, including Enterprise).
   - Classification: capability does not exist on any plan → real gap.

2. REAL GAP — Deal-D0D6B5
   - Prospect (HRIS Manager): "Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us."
   - Docs: HRIS integrations limited to Workday, BambooHR, Gusto, Rippling (Pro+); "ADP Workforce Now integrations are NOT currently listed as supported capabilities."
   - Classification: requested integration absent on every plan → real gap.

3. PLAN GATE — Deal-CFE7F4
   - Prospect (VP People): "I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?"
   - Docs: "Custom report builder" = Enterprise only (— on Core and Pro).
   - Classification: capability exists but is tier-gated above their current plan → plan gate, not a gap.

4. ROLLOUT/ENABLEMENT ISSUE — Deal-84DBA6
   - Prospect (People Ops Manager): "We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it."
   - Docs: "Slack and Microsoft Teams integration" = yes on all tiers.
   - Classification: supported feature, unused due to no training → enablement issue, not a gap.

5. NO PROSPECT-VOICED CANDIDATE — Deal-36C33F
   - The only gap statement (no native mobile app for hourly workers) was made by rep Alex Franklin, which per instructions does not count as prospect voice. The prospect (HR Manager) voiced no need: "Good to know. The web version should be fine for our office staff for now." Excluded.

**Arithmetic:** 5 transcripts → 4 prospect-voiced candidates → 2 real gaps + 1 plan gate + 1 enablement issue (1 excluded for rep-voiced).

**SUMMARY — true product gaps in active deals (2):**

1. SCIM user provisioning unavailable — Deal-EC3025 — blocking deal per prospect's security-approval requirement (amount not in data).
2. ADP Workforce Now integration unavailable — Deal-D0D6B5 — prospect-stated dealbreaker (amount not in data).
rep-performance 5 tests
stale-pipeline-by-repPASS 1.00
Plain-English question: Which open deals have had no email, call, or meeting in the last 7 days?
Prompt — what the model saw
The snapshot date is 2026-09-05. Using the open deals and the engagements-by-deal table, list every open deal with no logged email, call, or meeting in the last 7 days. Compute recency from the engagements table's last_email/last_call/last_meeting fields, NOT from any date field on the deal (the deal's last_contacted_field is unreliable). For each stale deal give: deal alias, owner name, stage, amount, days since last contact. Group by owner, order by amount descending within each owner. Then state, per owner, how many stale deals they carry and the total stale amount. (Do not write outreach text.)
cost $0.0949 · 291s · in 40,236 / out 59,927 tok
✓ stale_recall_ge_80✓ grouped_by_owner✓ jake_top_amount✓ kevin_stale_amt✓ jenn_stale_amt✓ stale_count

✓ Correct Answer (ground truth)

Plain-English answer: 65 open deals are stale; the largest is Deal-2D1F1B at $240,000, with Bryce Harmon carrying the most stale dollars at $626,243.
The answer in one line: jake_amt = 626,243; jenn_amt = 261,645; kevin_amt = 252,905.03; top_amount = 240,000; stale_count = 65; top_alias = Deal-2D1F1B
jake_amt626,243
jenn_amt261,645
kevin_amt252,905.03
stale_count65
top_amount240,000
top_aliasDeal-2D1F1B
graded checksstale_recall_ge_80 · grouped_by_owner · jake_top_amount · kevin_stale_amt · jenn_stale_amt · stale_count

Model's actual answer

show the model's full answer
**Method.** Stale = most recent of last_email / last_call / last_meeting (engagements table) older than 2026-08-29, i.e. > 7 days before the 2026-09-05 snapshot. Dates after the snapshot (future-dated last_meeting values, e.g. 2026-09-14) were excluded — a meeting dated after 2026-09-05 has not occurred; see caveat below. 156 open deals: 81 fresh, 75 stale, 2 with no engagement row. Owner order below is by total stale amount; deals within each owner are amount-descending.

**Bryce Harmon (119337721)** — 18 stale deals, $692,964 total
```
Deal-2D1F1B  DS1  240,000   81d (2026-06-16)
Deal-66D1FC  DS1   99,000   16d (2026-08-20)
Deal-950043  DS1   70,000   19d (2026-08-17)
Deal-B23205  DS1   45,000   16d (2026-08-20)
Deal-7BBDFA  DS3   37,440   46d (2026-07-21)
Deal-332637  DS2   36,000    9d (2026-08-27)
Deal-1BEEBF  DS1   31,500   19d (2026-08-17)
Deal-A414F6  DS1   25,200   19d (2026-08-17)
Deal-C5658B  DS1   23,400   16d (2026-08-20)
Deal-40522D  DS3   21,000   19d (2026-08-17)
Deal-C1FA6D  DS1   18,000   16d (2026-08-20)
Deal-01E193  DS1   12,600    8d (2026-08-28)
Deal-F0EBBB  DS3   11,400   24d (2026-08-12)
Deal-927338  DS1   10,920   18d (2026-08-18)
Deal-E25A09  DS1    6,000    9d (2026-08-27)
Deal-C9C286  DS2    5,502    9d (2026-08-27)
Deal-012CB1  DS1        1   23d (2026-08-13)
Deal-3795AD  DS2        1    8d (2026-08-28)
```

**Dana Mercer (83155923)** — 16 stale deals, $279,495 total
```
Deal-44EA29  DS2  60,000   10d (2026-08-26)
Deal-E51FB7  DS2  43,875   12d (2026-08-24)
Deal-B42F46  DS1  27,000   19d (2026-08-17)
Deal-BA3DDC  DS3  23,400   15d (2026-08-21)
Deal-9DDE86  DS2  20,000   15d (2026-08-21)
Deal-215CCA  DS3  18,900   17d (2026-08-19)
Deal-5EED42  DS3  16,250   11d (2026-08-25)
Deal-57887A  DS2  15,000    8d (2026-08-28)
Deal-944310  DS4  10,500   33d (2026-08-03)
Deal-B7EBD1  DS5   9,000   16d (2026-08-20)
Deal-3974EB  DS4   9,000    8d (2026-08-28)
Deal-F40F04  DS2   8,100   15d (2026-08-21)
Deal-7599B8  DS3   7,350   18d (2026-08-18)
Deal-87DDD1  DS1   5,000   19d (2026-08-17)
Deal-F336B6  DS3   4,200   15d (2026-08-21)
Deal-0660B4  DS4   1,920   16d (2026-08-20)
```

**Cole Ingram (83155924)** — 18 stale deals, $252,905.03 total
```
Deal-D04904  DS2  58,529.25  11d (2026-08-25)
Deal-B25F40  DS3  40,000      8d (2026-08-28)
Deal-813836  DS2  32,175     11d (2026-08-25)
Deal-1BA595  DS2  31,750     11d (2026-08-25)
Deal-CFE1E8  DS3  18,000     11d (2026-08-25)
Deal-CD47A6  DS2  12,168     11d (2026-08-25)
Deal-627646  DS3  11,193     11d (2026-08-25)
Deal-FF809F  DS2   7,781.20  11d (2026-08-25)
Deal-AF932D  DS2   7,225.40  11d (2026-08-25)
Deal-A71728  DS2   6,947.50  11d (2026-08-25)
Deal-8BC9F5  DS2   5,616     10d (2026-08-26)
Deal-175395  DS3   4,779.88  11d (2026-08-25)
Deal-481E24  DS3   4,140     10d (2026-08-26)
Deal-C7F9BF  DS2   3,360     11d (2026-08-25)
Deal-2F3A66  DS3   3,334.80  11d (2026-08-25)
Deal-342E96  DS2   2,700     24d (2026-08-12)
Deal-E568D5  DS3   1,875     11d (2026-08-25)
Deal-FD9F4E  DS5   1,330     10d (2026-08-26)
```

**Alex Franklin (84342457)** — 20 stale deals, $113,936 total
```
Deal-CC08D1  DS1  24,000  16d (2026-08-20)
Deal-E73427  DS3  18,000  10d (2026-08-26)
Deal-885F45  DS2   9,300  12d (2026-08-24)
Deal-C2FF3C  DS1   8,316  10d (2026-08-26)
Deal-3EED2C  DS2   7,200  n/a — no engagement row (last contact unknown)
Deal-0D2F7A  DS3   5,100  12d (2026-08-24)
Deal-6C60D4  DS3   4,800  12d (2026-08-24)
Deal-13FEBD  DS2   4,680  12d (2026-08-24)
Deal-819506  DS1   4,400   8d (2026-08-28)
Deal-9D0060  DS3   3,840  12d (2026-08-24)
Deal-690476  DS2   3,600  18d (2026-08-18)
Deal-C6D97A  DS4   3,240   8d (2026-08-28)
Deal-EE195F  DS3   3,120   8d (2026-08-28)
Deal-278DEC  DS3   2,700   8d (2026-08-28)
Deal-635B8E  DS3   2,600  18d (2026-08-18)
Deal-6883F3  DS1   2,400  16d (2026-08-20)
Deal-4A13AD  DS3   2,160  26d (2026-08-10)
Deal-F67D31  DS2   1,800   8d (2026-08-28)
Deal-5FDCE4  DS3   1,600  12d (2026-08-24)
Deal-BA571A  DS4   1,080  18d (2026-08-18)
```

**Farid Osman (716654662)** — 2 stale deals, $2,881 total
```
Deal-8BA24E  DS1  2,880   8d (2026-08-28)
Deal-8FDCD2  DS1      1  15d (2026-08-21)
```

**Elena Sinclair (701163055)** — 1 stale deal, $2,100 total
```
Deal-57FF13  DS1  2,100  n/a — no engagement row (last contact unknown)
```

**Per-owner stale summary**
| Owner | Stale deals | Total stale amount |
|---|---|---|
| Bryce Harmon | 18 | $692,964 |
| Dana Mercer | 16 | $279,495 |
| Cole Ingram | 18 | $252,905.03 |
| Alex Franklin | 20 | $113,936 |
| Farid Osman | 2 | $2,881 |
| Elena Sinclair | 1 | $2,100 |
| **Total** | **75** | **$1,344,281.03** |

**Caveats (explicit).**
1. Two deals (Deal-3EED2C, Deal-57FF13) have no row in the engagements table, so days-since-last-contact cannot be computed; they are listed as stale on the basis of no logged contact in the 90-day window.
2. 17 deals carry last_meeting dates after the snapshot (future meetings, e.g. 2026-09-14); these were not counted as contact. Ten of those are in the stale list above (Deal-944310, Deal-3795AD, Deal-01E193, Deal-C1FA6D, Deal-927338, Deal-A414F6, Deal-819506, Deal-7599B8, Deal-8BA24E, Deal-8FDCD2). If you instead treat a future-dated last_meeting as contact, those 10 drop out and the stale count falls to 65 (Bryce 15/$652,563, Dana 15/$268,995, Cole 18/$252,905.03, Alex 17/$95,416, Farid 1/$2,880, Elena 1/$2,100).
activity-mix-vs-outcomePASS 1.00
Plain-English question: Per rep, what did the last 30 days of emails, calls, and meetings look like next to the deals that advanced?
Prompt — what the model saw
The snapshot date is 2026-09-05. For each sales rep (owner), compute the last-30-day totals of emails, calls, and meetings from the engagements table, and the number of deals that entered DS2 in the last 30 days (t_ds2 within 30 days of the snapshot). Report per rep: the activity mix as percentages (emails/calls/meetings share of that rep's total activities), total activities per DS2 entry, and rank the reps by that efficiency ratio (lowest activities per DS2 entry = most efficient). Name the most efficient rep and the highest-volume rep, and note if they differ. Do not attribute cause.
cost $0.1097 · 317s · in 76,749 / out 63,689 tok
✓ alex_ds2_30d✓ justin_volume_high✓ efficient_named✓ mix_present✓ no_causal_claim

✓ Correct Answer (ground truth)

Plain-English answer: Alex Franklin is the standout signal with 18 deals entering DS2 in 30 days (next closest: Bryce Harmon with 4), alongside per-rep activity totals.
The answer in one line: alex_ds2_30d = 18; bryce_ds2_30d = 4; note = legacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
alex_ds2_30d18
bryce_ds2_30d4
notelegacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
graded checksalex_ds2_30d · justin_volume_high · efficient_named · mix_present · no_causal_claim

Model's actual answer

show the model's full answer
Snapshot 2026-09-05. Window for "last 30 days" = 2026-08-06 through 2026-09-05 (inclusive). Activities = emails_30d + calls_30d + meetings_30d summed per rep across all open deals in deals_open.csv, joined to engagements_by_deal_90d.csv by deal_id. DS2 entries = distinct deals with t_ds2 in the window.

**Per-rep results (ranked by efficiency, lowest activities per DS2 entry first)**

| Rep | Emails | Calls | Meetings | Total | Mix E/C/M | DS2 entries | Acts per DS2 |
|---|---|---|---|---|---|---|---|
| 1. Alex Franklin | 307 | 36 | 41 | 384 | 79.9 / 9.4 / 10.7% | 18 | 384/18 = **21.33** |
| 2. Bryce Harmon | 162 | 0 | 43 | 205 | 79.0 / 0.0 / 21.0% | 4 | 205/4 = **51.25** |
| 3. Cole Ingram | 96 | 14 | 1 | 111 | 86.5 / 12.6 / 0.9% | 2 | 111/2 = **55.50** |
| 4. Farid Osman | 38 | 0 | 34 | 72 | 52.8 / 0.0 / 47.2% | 1 | 72/1 = **72.00** |
| 5. Dana Mercer | 84 | 18 | 11 | 113 | 74.3 / 15.9 / 9.7% | 1 | 113/1 = **113.00** |

DS2-entering deal aliases by rep:
- Alex Franklin (18): Deal-EE195F, Deal-D9A72E, Deal-7FA0C3, Deal-E531A6, Deal-36C33F, Deal-D1E6C2, Deal-317E6F, Deal-4F775F, Deal-F436DA, Deal-CA5E44, Deal-46988D, Deal-898FC5, Deal-E73427, Deal-5296C9, Deal-92D97D, Deal-403845, Deal-3EED2C, Deal-1FC049
- Bryce Harmon (4): Deal-25F752, Deal-CA7DC0, Deal-D73B89, Deal-1CCE5C
- Cole Ingram (2): Deal-1BA595, Deal-42326B
- Dana Mercer (1): Deal-57887A
- Farid Osman (1): Deal-499BF6

**Most efficient:** Alex Franklin (21.33 activities per DS2 entry).
**Highest volume:** Alex Franklin (384 total 30d activities). They are the same rep — most efficient and highest volume do not differ.

**Data notes (explicit):**
- Elena Sinclair: no computable metrics. Her only open deal (Deal-57FF13) has no row in engagements_by_deal_90d.csv and no t_ds2, so she has 0 DS2 entries and unknown activity (no ratio; excluded from ranking, not scored as 0).
- Two deals have no engagement row, so their 30d activities are unknown and excluded from sums: Deal-3EED2C (Alex Franklin) and Deal-57FF13 (Elena Sinclair). Alex's totals are therefore possibly understated.
- Archived owners Gavin Porter and Hugo Lindqvist own no open deals in the file (no data).
- Boundary sensitivity: three entries sit exactly 30 days out (2026-08-06): Deal-1CCE5C (Bryce), Deal-EE195F and Deal-D9A72E (Alex). Excluding them leaves Alex most efficient (24.00); order below Alex becomes Cole 55.50, Farid 72.00, Dana 113.00, Bryce 68.33.
- Arithmetic shown in the table; mix percentages are each channel ÷ that rep's total (e.g., Alex emails 307/384 = 79.9%).

No cause attribution made; totals are correlational counts only.
ae-qtd-scorecardPASS 1.00
Plain-English question: Build Alex Franklin's QTD scorecard: bookings versus quota, pipeline, and conversion.
Prompt — what the model saw
Build a QTD scorecard for the named AE as of 2026-09-05: bookings vs quota with attainment percent, new vs expansion split, active pipeline by stage with amount, rolling 90-day DS2-to-won rate, win and loss counts with top loss reason, activity volume by type in the last 30 days. Exclude any closed-won deal dated before the quarter. Close with three coaching observations grounded in the numbers.
cost $0.0731 · 266s · in 21,297 / out 48,767 tok
✓ bookings✓ quota✓ attainment✓ new_expansion_split✓ top_loss_reason✓ coaching_obs

✓ Correct Answer (ground truth)

Plain-English answer: 75% attainment ($150,000 on a $200,000 quota), $113,500 new versus $36,500 expansion, pipeline by stage, and the rolling DS2-to-won rate.
The answer in one line: quota = 200000; bookings = 150000; new_bookings = 113500; expansion_bookings = 36500; open_deals_count = 125; attainment_pct = 75; owner = Alex Franklin (+12 more below)
quota200,000
attainment_pct75
bookings150,000
expansion_bookings36,500
expansion_count3
losses_qtd27
new_bookings113,500
new_count5
open_deals_count125
wins_qtd8
activity_30d_totalsemails_30d=807; calls_30d=112; meetings_30d=128; notes_30d=50
allowed_new18 items — e.g. Deal-A1C3E5, Deal-B7D2F4, Deal-C9E1A6, Deal-D4B8C2
ds2_to_wonwindow_start=2026-06-07; won_in_window=8; closed_in_window=35; rate=0.2286
ownerAlex Franklin
pipeline_by_stage{"DS1": {"count": 20, "amount": 284621}, "DS2": {"count": 28, "amount": 353760}, "DS3": {"count": 67, "amount": 552705}, "DS4": {"count": 5, "amount": 23574}, "DS5": {"count": 5, "amount": 45730}}
prior_quarter_excludeddeal_alias=Deal-B3E6F1; close_date=2026-06-20; amount=24000
quarter2026-Q3
snapshot_date2026-09-05
top_loss_reasonreason=Lost- Timing (1 year or more); count=13
graded checksbookings · quota · attainment · new_expansion_split · top_loss_reason · coaching_obs

Model's actual answer

show the model's full answer
AE SCORECARD — Alex Franklin, 2026-Q3, as of snapshot 2026-09-05
Sources: ae_deals.csv (161 rows), ae_engagements.csv, quota.csv. Quarter = 2026-07-01 to 2026-09-05.

1) BOOKINGS VS QUOTA
QTD closed-won: 8 deals
  $40,000 + $20,000 + $35,000 + $21,000 + $12,000 + $11,000 + $6,500 + $4,500 = $150,000
Attainment: $150,000 / $200,000 = 75.0%. Gap to quota: $50,000.
Excluded per instruction: Deal-B3E6F1, $24,000, closed-won 2026-06-20 (pre-quarter). No other QTD closed-won existed before 07-01.

2) NEW VS EXPANSION (QTD)
New: 5 deals, $113,500 (75.7%) — Deal-A1C3E5 $40k, Deal-B7D2F4 $35k, Deal-C9E1A6 $21k, Deal-D4B8C2 $11k, Deal-E6F3A9 $6.5k
Expansion: 3 deals, $36,500 (24.3%) — Deal-F2C7D8 $20k, Deal-A8B4D6 $12k, Deal-C5D9E2 $4.5k

3) ACTIVE PIPELINE BY STAGE (status=open, 125 deals)
  DS1: 20 deals, $284,621 (22.6%)
  DS2: 28 deals, $353,760 (28.1%)
  DS3: 67 deals, $552,705 (43.9%)
  DS4:  5 deals, $23,574 (1.9%)
  DS5:  5 deals, $45,730 (3.6%)
  Total: $1,260,390
Note: Deal-7A2454 (DS3, $1,275) shows close_date 2026-09-04 but is still open — already slipped past its date. Of the $1.26M, only $109,363 (22 deals) is dated to close in September; $10,920 of that sits in DS4/DS5.

4) ROLLING 90-DAY DS2-TO-WON (window 2026-06-07 to 2026-09-05, by entered_ds2)
23 open deals currently in DS2 entered DS2 within the window; 0 have converted to won → 0/23 = 0.0%.
(The 28 total DS2 deals include 5 entered before the window, e.g. Deal-5BFE3B entered 2026-01-05.) Cross-check via closed outcomes whose DS2 entry falls in-window: 8 won / 35 closed = 22.9% — the stage itself wins at ~1-in-4.5, but the current open DS2 cohort has produced nothing.

5) WINS / LOSSES QTD
Won: 8. Lost: 27. Close rate: 8/35 = 22.9%. Lost value $329,272 = 2.2x QTD bookings.
Top loss reason: "Lost- Timing (1 year or more)" — 13/27 losses (48.1%), $184,681 of lost value (56.1%). Next: MIA 5, Competitor 5, Lost DM 2, Feature Request 1, Does-not-fit-ICP 1. Avg lost deal $12,195 vs avg won deal $18,750.

6) ACTIVITY, LAST 30 DAYS (per ae_engagements.csv)
Emails 807 (73.6%), Meetings 128 (11.7%), Calls 112 (10.2%), Notes 50 (4.6%). Total 1,097.
Compare: the 9 won deals averaged 11.0 emails + 2.8 meetings/deal; the 125 open deals average 4.7 emails + 0.7 meetings/deal. 82 of 125 open deals (65.6%) had zero meetings in 30 days; Deal-3EED2C had zero activity of any type.

COACHING OBSERVATIONS
1. Timeline qualification is the leak, not effort. 13 of 27 losses (48%) and $184,681 (56%) of lost value were "Timing (1 year or more)" — while activity skews 73.6% email. Won deals carried ~2.8 meetings each vs 0.7 on open deals; pushing live-conversation timeline qualification before deals reach DS2/DS3 would catch the 12-month-out buyers earlier.
2. DS2 is where pipeline goes to stall. 0/23 of the 90-day DS2 cohort has won, and 28 deals / $353,760 now sit in DS2, some entered as far back as January (Deal-5BFE3B). Meanwhile the stage's historical close rate is 22.9% — the deals aren't bad, they're parked. A DS2 aging/slip review is the single highest-leverage cleanup.
3. The $50k gap is coverable but thin. Only $109,363 of the $1.26M pipeline is dated to close in September, and just $10,920 of that is in late stages (DS4/DS5: $69,304 total = 5.5% of pipe). With 82 open deals untouched by a meeting in 30 days, the fastest path to the gap is re-engaging the stale DS2/DS3 deals with meetings, not adding net-new.
multithreading-gapPASS 1.00
Plain-English question: Which open deals are single-threaded or under-threaded?
Prompt — what the model saw
Find every open deal that is single-threaded (fewer than 2 active contacts) or under-threaded (fewer than 3, or all contacts in one persona). Active means engaged in the last 60 days and not marked former. Personas: economic buyer, champion, HR admin, IT security, finance. For each flagged deal: amount, stage, active contact count, personas present, personas missing, the most valuable persona to add given the stage, and any on-file unengaged contact at that company who fits; otherwise say none on file.
cost $0.0176 · 72s · in 6,186 / out 9,371 tok
✓ flagged_recall✓ persona_language✓ single_thread_term✓ on_file_fit

✓ Correct Answer (ground truth)

Plain-English answer: Deals with fewer than 2 active contacts (single-threaded) or fewer than 3, or all in one persona (under-threaded), each flagged with the missing persona.
The answer in one line: active_window_start = 2026-07-07
active_window_start2026-07-07
allowed_new59 items — e.g. C-077A0E, C-0D15DF, C-1BDFE0, C-2AE3AA
best_persona_ruleDS2/DS3: economic buyer > IT security > finance > HR admin > champion; DS4/DS5: economic buyer > finance > IT security > HR admin > champion; first persona not present among active contacts
flag_ruleflag if active_contacts < 2 (single-threaded) or active_contacts < 3 or all active contacts share one persona (under-threaded); active = engaged within 60 days and not former
flaggeddeal_alias=Deal-36C33F; amount=15000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-50D386; amount=36000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5408B0; amount=14850; stage=DS4; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5BFE3B; amount=18000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-885F45; amount=9300; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=IT security | deal_alias=Deal-92D97D; amount=60000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer … (+5 more)
healthy_not_flaggedDeal-4B0BEB, Deal-84DBA6, Deal-D348E1
persona_universeeconomic buyer, champion, HR admin, IT security, finance
graded checksflagged_recall · persona_language · single_thread_term · on_file_fit

Model's actual answer

show the model's full answer
Note on missing fields: neither file contains deal amount or stage. Those are stated as missing below; persona recommendations are therefore coverage-based, not stage-based. Activity window: as-of 2026-09-25 minus 60 days = active means last_engaged_date ≥ 2026-07-27 and is_former = false.

FLAGGED DEALS (10 of 14; amount/stage not on file)

1) Deal-EC3025 (C-FDD0C7) — amount: not on file; stage: not on file
- Active contacts: 1 (CT-047C54, champion, eng 2026-09-02). CT-F2C1AE (economic buyer) excluded: is_former=true.
- Personas present: champion | Missing: economic buyer, HR admin, IT security, finance
- Flags: single-threaded; single-persona. Prior EB is former → highest risk.
- Most valuable add: economic buyer. On-file fit: CT-6827DB, Chief People Officer (unengaged list).

2) Deal-92D97D (C-E23238) — amount/stage: not on file
- Active contacts: 1 (CT-01F5B4, HR admin, eng 2026-08-28). CT-A902AE (champion) stale: eng 2026-06-01 = 116 days (>60).
- Personas present: HR admin | Missing: economic buyer, champion, IT security, finance
- Flags: single-threaded; single-persona.
- Most valuable add: economic buyer. On-file fit: none on file. (Champion also missing with no on-file match; CT-A902AE exists but is stale 116 days.)

3) Deal-50D386 (C-EB10E4) — amount/stage: not on file
- Active contacts: 2 (CT-AA41B2 champion 2026-09-01; CT-B9C35B HR admin 2026-08-25).
- Personas present: champion, HR admin | Missing: economic buyer, IT security, finance
- Flags: under-threaded (<3 active).
- Most valuable add: economic buyer. On-file fit: CT-A1C4B3, Chief People Officer.

4) Deal-D0D6B5 (C-32918E) — amount/stage: not on file
- Active contacts: 3 (CT-87CED4 2026-09-02; CT-DE6D7C 2026-08-19; CT-FD70B2 2026-08-07) — all champion.
- Personas present: champion only | Missing: economic buyer, HR admin, IT security, finance
- Flags: single-persona (meets the ≥3 count but zero persona diversity).
- Most valuable add: economic buyer. On-file fit: CT-1FA4DB, Chief People Officer.

5) Deal-5BFE3B (C-535D36) — amount/stage: not on file
- Active contacts: 2 (CT-57123B 2026-08-31; CT-5CE757 2026-08-12) — both champion.
- Personas present: champion only | Missing: economic buyer, HR admin, IT security, finance
- Flags: under-threaded (<3); single-persona.
- Most valuable add: economic buyer. On-file fit: none on file.

6) Deal-36C33F (C-077A0E) — amount/stage: not on file
- Active contacts: 1 (CT-4FE556, IT security, eng 2026-08-15). CT-405B45 (champion) and CT-86B22F (economic buyer) both is_former=true.
- Personas present: IT security | Missing: economic buyer, champion, HR admin, finance
- Flags: single-threaded; single-persona. Champion AND economic buyer both lost → highest severity in the set.
- Most valuable add: economic buyer. On-file fit: CT-1DB73E, Chief People Officer. (Champion: none on file.)

7) Deal-885F45 (C-5E8EFB) — amount/stage: not on file
- Active contacts: 2 (CT-51C81E economic buyer 2026-08-26; CT-D9A0E8 champion 2026-08-11).
- Personas present: economic buyer, champion | Missing: HR admin, IT security, finance
- Flags: under-threaded (<3 active).
- Most valuable add: IT security (EB + champion already active; IT security is both missing and the on-file fit). On-file fit: CT-B3F25D, IT Security Lead.

8) Deal-FCBE5B (C-737030) — amount/stage: not on file
- Active contacts: 1 (CT-4A5317, champion, eng 2026-08-29). No other contacts on the deal.
- Personas present: champion | Missing: economic buyer, HR admin, IT security, finance
- Flags: single-threaded; single-persona.
- Most valuable add: economic buyer. On-file fit: none on file.

9) Deal-5408B0 (C-2AE3AA) — amount/stage: not on file
- Active contacts: 2 (CT-D33AE4 champion 2026-09-01; CT-8742FD HR admin 2026-08-18).
- Personas present: champion, HR admin | Missing: economic buyer, IT security, finance
- Flags: under-threaded (<3 active).
- Most valuable add: economic buyer. On-file fit: CT-07FA76, Chief People Officer.

10) Deal-C6D97A (C-5A8FC2) — amount/stage: not on file
- Active contacts: 3 (CT-223DDC 2026-08-31; CT-B03555 2026-08-20; CT-4E8A2B 2026-08-05) — all champion.
- Personas present: champion only | Missing: economic buyer, HR admin, IT security, finance
- Flags: single-persona.
- Most valuable add: economic buyer. On-file fit: none on file.

11) Deal-F9A08A (C-0D15DF) — amount/stage: not on file
- Active contacts: 1 (CT-931B10, champion, eng 2026-09-03). CT-913581 (economic buyer) excluded: eng 2026-06-20 = 97 days (>60), not former but inactive.
- Personas present: champion | Missing: economic buyer, HR admin, IT security, finance
- Flags: single-threaded; single-persona.
- Most valuable add: economic buyer. On-file fit: CT-697541, Chief People Officer. (Separately: CT-913581 is on the deal but 97 days stale — a re-engagement target in addition to the on-file contact.)

NOT FLAGGED (3 of 14)
- Deal-84DBA6 (C-BE9BE0): 3 active, personas = economic buyer, champion, IT security. Missing finance, HR admin — meets thresholds.
- Deal-4B0BEB (C-1BDFE0): 4 active, personas = economic buyer, champion, HR admin, finance. Missing IT security — meets thresholds.
- Deal-D348E1 (C-804F29): 5 active, all 5 personas present.

Summary counts: 14 deals total → 10 flagged (5 single-threaded, 7 with <3 active, 7 single-persona — flags overlap). 7 of 10 flagged deals have a matching unengaged on-file contact; 3 have none on file (Deal-92D97D, Deal-5BFE3B, Deal-FCBE5B; plus no champion fit for Deal-36C33F and no coverage fit for Deal-C6D97A). Recommendation logic caveat: "most valuable persona given the stage" could not be stage-weighted because stage is not in the provided data; economic buyer was ranked first wherever no active EB exists, since all 10 flagged deals lack an active economic buyer.
talk-track-from-transcriptsPASS 1.00
Plain-English question: From this rep's last 10 calls: how do they open, handle objections, and how often do they lock a next step?
Prompt — what the model saw
From this rep's last 10 calls: what they lead with in the first five minutes, how they handle the three most common objections, how often a concrete next step is agreed (rate), and every competitor a prospect raised. Quote at most one line per finding and only lines that appear in the transcripts. End with two coaching notes.
cost $0.0064 · 26s · in 2,745 / out 2,947 tok
✓ objection_count✓ next_step_rate✓ competitors✓ coaching✓ quote_used

✓ Correct Answer (ground truth)

Plain-English answer: 8 of 10 calls open with the same customer-win story; objection handling, next-step agreement rate, and every competitor raised, with quotes.
The answer in one line: n_transcripts = 10; objection_count = 3; rep = Alex Franklin (+3 more below)
n_transcripts10
objection_count3
allowed_new13 items — e.g. TT-001, TT-002, TT-003, TT-004
competitor_quotesAwardco=We're also in late talks with Awardco — their rewards catalog looks bigger than yours.; Kudos=How are you different from Kudos? Our CEO used them at her last company.
competitors_prospect_raisedAwardco, Kudos
next_stepagreed=7; total=10; rate=0.7
objectionsobjection=budget; count=4; quote=Honestly, budget is locked until next fiscal year — I can't add a new line item right now.; transcript_id=TT-001 | objection=timing; count=3; quote=This is good, but can we revisit it next quarter? Open enrollment starts in October and we're underwater.; transcript_id=TT-002 | objection=in-house; count=3; quote=We already do recognition with a spreadsheet and quarterly gift cards — why would we change?; transcript_id=TT-004
opener_themetheme=opens with a customer recognition win story (400-person retailer, 18% turnover reduction); count=8; total=10; quote=Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it.; transcript_id=TT-001
repAlex Franklin
rep_raised_competitors_excludedWorkhuman
graded checksobjection_count · next_step_rate · competitors · coaching · quote_used

Model's actual answer

show the model's full answer
ANALYSIS: Alex Franklin — last 10 calls (TT-001 to TT-010, 2026-07-28 to 2026-09-03)

1) WHAT HE LEADS WITH (first five minutes)
- 8/10 calls open with the same customer-story opener: TT-001, -002, -003, -005, -006, -007, -008, -010 (verbatim identical).
  "Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it."
- 1/10 agenda-first (Deal-403845 / TT-004): "I put together a short agenda — security review first, then pricing."
- 1/10 pricing-first callback (Deal-1E2498 / TT-009): "You asked for straight pricing last time, so let's start there."
- In one of the 8 (Deal-C61CF7 / TT-005), he adds a competitor jab unprompted at minute 2: "And unlike Workhuman, our pricing includes the full rewards catalog with no extra margin."

2) THREE MOST COMMON OBJECTIONS + HIS HANDLING
Arithmetic: prospect objection lines by type — budget-locked 4 (TT-001, -003, -006, -010), timing/open-enrollment 3 (TT-002, -005, -008), spreadsheet status quo 3 (TT-004, -007, -009), committee/no-commit 2 (TT-004, -010), no urgency 1 (TT-007). Total objection lines = 13.

a. "Budget is locked until next fiscal year" — 4/10. Handling (used 4/4, verbatim identical): turnover-savings ROI reframe.
   "Totally fair. Most teams fund this out of turnover savings — that retailer saved about $210k in avoided backfills, which is how their finance team signed off."

b. "Revisit next quarter / open enrollment" — 3/10. Handling (3/3, verbatim identical): 90-day pilot reframe.
   "Makes sense. What if we scope a 90-day pilot with one department so you have internal data before next quarter's planning?"

c. "We already do recognition with a spreadsheet and quarterly gift cards" — 3/10. Handling (3/3, verbatim identical): scale/automation contrast.
   "Spreadsheets work until they scale — the difference is automation: milestones fire without HR lifting a finger, and you get analytics on who is being recognized."

Unscripted cases: on committee waits he has no play — concedes ("Understood — I'll leave it with you." TT-004; "Understood, thanks for the candor." TT-010). On "no urgency" he concedes with "Fair enough." (TT-007).

3) NEXT-STEP AGREEMENT RATE
His ask, verbatim in 7 calls: "Should we lock the next step — a working session with your team this week?"
Prospect agreement: "Yes, Thursday at 2pm works — send the invite and I'll bring our HRIS manager."
Arithmetic: agreed = TT-001, -002, -003, -005, -006, -008, -009 = 7; not agreed = TT-004, -007, -010 = 3. Rate = 7/10 = 70%.

4) COMPETITORS RAISED BY PROSPECTS (exhaustive)
- Awardco — Deal-547B2B (TT-003): "We're also in late talks with Awardco — their rewards catalog looks bigger than yours."
- Kudos — Deal-EDC141 (TT-007): "How are you different from Kudos? Our CEO used them at her last company."
Data-fidelity note: Workhuman (Deal-C61CF7, TT-005) appears only in the rep's own line, never raised by a prospect, so it is excluded from the prospect list above.

COACHING NOTES
1. Next-step rate collapses exactly when the blocker is internal process: 7/10 overall but 0/3 on committee/no-commit/no-urgency calls (Deal-403845, Deal-EDC141, Deal-84DBA6). He has no counter-move for authority stalls — "Understood — I'll leave it with you" ends the call with nothing booked. Coach a fallback step (offer to join the budget-committee meeting, multithread a champion, or set a dated check-in) so a stall never ends in no next step.
2. The playbook is fully canned: opener identical in 8/10 and all three objection responses verbatim-repeated, and it fails as a stack — Deal-EDC141 ran competitor (Kudos) + status quo + "I need to think about it — there's no urgency on our side" and ended in "Fair enough." Add a discovery/urgency play for that pattern, and cut the unprompted competitor mention (Workhuman, Deal-C61CF7) — it introduces a name the prospect never raised.
reporting-analytics 5 tests
q3-weighted-forecastPASS 1.00
Plain-English question: Build the Q3 2026 weighted forecast from the open pipeline, and show the math deal by deal.
Prompt — what the model saw
Below are the open deals in the default sales pipeline with a close date between 2026-07-01 and 2026-10-15, pulled from the CRM warehouse on 2026-09-05. Stages: DS1 (discovery) through DS5 (contract). Build the Q3 2026 forecast (quarter = 2026-07-01 to 2026-09-30). Weighting rule: 100 percent of COMMIT plus 35 percent of BEST_CASE. PIPELINE deals count zero. Only deals with a close date inside the quarter count. Report: COMMIT total, BEST_CASE total, weighted forecast, count of deals in each category inside the quarter, the deals you excluded for being outside the quarter (count and total amount), and the top 5 BEST_CASE deals by amount inside the quarter. Show arithmetic. Then, in a final '## Data quality' section of no more than four sentences, name the data quality problems in this extract that would make the forecast unreliable if it were run unattended.
cost $0.0351 · 108s · in 28,045 / out 16,531 tok
✓ commit_total_44729✓ best_case_total_203565✓ weighted_115976.75✓ commit_count_7✓ best_case_count_24✓ excluded_count_32✓ excluded_amount_227575✓ top5_aliases✓ flags_blank_owner✓ flags_missing_why_buys

✓ Correct Answer (ground truth)

Plain-English answer: A stage-weighted Q3 forecast number computed from the open pipeline, with per-stage math that reconciles to the deal data.
The answer in one line: excluded_amount = 227575; best_case_total = 203565; weighted_forecast = 115,976.75; commit_total = 44729; excluded_out_of_quarter = 32; best_case_count = 24; flags = blank owner column + empty why_buys fields must be surfaced (+2 more below)
excluded_amount227,575
best_case_count24
best_case_total203,565
commit_count7
commit_total44,729
excluded_out_of_quarter32
weighted_forecast115,976.75
flagsblank owner column + empty why_buys fields must be surfaced
top5_best_caseDeal-2D7423, Deal-25F752, Deal-E53952, Deal-5EED42, Deal-FA32A0
graded checkscommit_total_44729 · best_case_total_203565 · weighted_115976.75 · commit_count_7 · best_case_count_24 · excluded_count_32 · excluded_amount_227575 · top5_aliases · flags_blank_owner · flags_missing_why_buys

Model's actual answer

show the model's full answer
# Q3 2026 Forecast (quarter 2026-07-01 → 2026-09-30; weighted = COMMIT×100% + BEST_CASE×35%)

**COMMIT total (in-quarter): $44,729.00 — 7 deals**
- Deal-547B2B 11,200 (09-11)
- Deal-B7EBD1 9,000 (09-10)
- Deal-403845 9,000 (09-11)
- Deal-A2B47C 6,360 (09-11)
- Deal-2465CE 5,400 (09-10)
- Deal-A5E80A 2,520 (09-11)
- Deal-499BF6 1,249 (09-30)

Arithmetic: 11,200 + 9,000 + 9,000 + 6,360 + 5,400 + 2,520 + 1,249 = **44,729.00**

**BEST_CASE total (in-quarter): $203,565.00 — 24 deals**
Arithmetic: 38,935 + 24,000 + 19,656 + 16,250 + 11,116 + 10,800 + 10,500 + 9,890 + 9,720 + 9,000 + 7,200 + 3,840 + 3,780 + 3,600 + 3,240 + 3,150 + 3,120 + 3,060 + 2,916 + 2,760 + 2,484 + 2,100 + 1,920 + 528 = **203,565.00**

**PIPELINE (in-quarter): 23 deals, $201,637.40 — counts zero under the rule.**

**Weighted forecast**
44,729.00 + 0.35 × 203,565.00 = 44,729.00 + 71,247.75 = **$115,976.75**

**Deal counts inside the quarter: 54 total** — COMMIT 7, BEST_CASE 24, PIPELINE 23.

**Excluded for being outside the quarter: 32 deals, $227,575.00**
All 32 have close dates 2026-10-01 through 2026-10-15 (none before 2026-07-01). Caution: these include 1 COMMIT (Deal-D348E1, $13,770, 2026-10-15) and 9 BEST_CASE deals totaling $28,240 that a naive "close_date ≤ 2026-10-15" filter would have wrongly included. Largest excluded deal: Deal-E51FB7, $43,875, closing 2026-10-01 — one day outside the boundary.

**Top 5 BEST_CASE deals in-quarter**
1. Deal-2D7423 — $38,935 (DS3, 2026-09-30)
2. Deal-25F752 — $24,000 (DS4, 2026-09-25)
3. Deal-E53952 — $19,656 (DS4, 2026-09-30)
4. Deal-5EED42 — $16,250 (DS3, 2026-09-30)
5. Deal-FA32A0 — $11,116 (DS3, 2026-09-25)

## Data quality
Owner is blank on 85 of 86 rows (only Deal-C9C286 names Bryce Harmon), so the forecast cannot be validated or coached by rep. Category hygiene is broken: Deal-A5E80A (DS1) and Deal-499BF6 (DS2) carry COMMIT despite being pre-DS5, and only 2 of 31 in-quarter COMMIT/BEST deals have any "why buys" rationale (why_buys_chars > 0), meaning the weighted number rests largely on undocumented judgment. Four "open" deals (Deal-31AD2C, Deal-333EBB, Deal-57FF13, Deal-7A2454) have close dates already before the 2026-09-05 pull date, signaling unmanaged slippage, while boundary slippage risk is high with the largest deal (Deal-E51FB7, $43,875) sitting one day outside the quarter and Deal-D348E1 COMMIT just past it. Finally, Deal-42326B carries a fractional amount (2,480.4), suggesting proration or a data-entry artifact that should be confirmed before publishing.
aha-moment-2x2PASS 1.00
Plain-English question: Do companies that hit 5+ givers and one redemption in month one actually retain better at 24 months?
Prompt — what the model saw
Test the activation hypothesis on the company cohort below: companies with 5+ unique givers (m1_users >= 5) AND 1+ successful redemption (m1_redemptions >= 1) in their first calendar month retain better at 24 months than those with one signal or neither. Every company in this extract is already 25+ months old, so current_status = 'active' means retained at 24 months. Report the full 2x2 (cohort size and 24-month retention rate for: both signals, givers-only, redemption-only, neither), state how many companies were excluded from the denominator and why (if none, say so), name the single signal with the largest retention lift, and state what this does and does not prove.
cost $0.0406 · 108s · in 22,660 / out 16,745 tok
✓ both_66✓ givers_47✓ redemp_31✓ neither_40✓ excluded_0✓ largest_lift_givers

✓ Correct Answer (ground truth)

Plain-English answer: The four cohort cells are graded against computed 24-month retention; the numbers come from the scorer's cohort table, not opinion.
The answer in one line: both_5plus_givers_and_redemption = 66.0% (31 of 47)
both_5plus_givers_and_redemption66.0% (31 of 47)
exclusionsnone — claim bounded as correlation, not causation
givers_only46.9% (23 of 49)
largest_single_signal_liftunique givers
neither40.0% (38 of 95)
redemption_only31.0% (9 of 29)
graded checksboth_66 · givers_47 · redemp_31 · neither_40 · excluded_0 · largest_lift_givers

Model's actual answer

show the model's full answer
Cohort tested: 220 companies (signups 2023-01 through 2023-07), all classifiable.

**Method.** Givers signal = m1_users >= 5 (per your operationalization of unique givers — the extract has no true giver count, so m1_users stands in). Redemption signal = m1_redemptions >= 1. Retained = current_status = 'active'; both 'cancelled' (214) and 'non_renewing' (6) count as not retained.

**2x2 results (n = 220)**

| Cohort | n | Retained | 24-mo retention | Lift vs neither |
|---|---|---|---|---|
| Both signals | 47 | 31 | 31/47 = 65.96% | +25.96 pts |
| Givers-only | 49 | 23 | 23/49 = 46.94% | +6.94 pts |
| Redemption-only | 29 | 9 | 9/29 = 31.03% | −8.97 pts |
| Neither | 95 | 38 | 38/95 = 40.00% | baseline |

**Exclusions: none.** All 220 rows have integer m1_users/m1_redemptions and a status in {active, cancelled, non_renewing}; every company appears in exactly one cell (47+49+29+95 = 220). Nothing was dropped from the denominator.

**Largest single-signal lift: the givers signal.** Givers-only (+6.94 pts over neither) is the only signal with a positive lift on its own. Redemption-only is *negative* (−8.97 pts, 31.03% vs 40.00%) — a first-month redemption without 5+ givers is associated with worse retention than doing nothing measurable. The combination is what carries the effect: 65.96% vs givers-only's 46.94% (+19.02 pts from adding redemption on top of givers) vs +12.02 pts from adding givers on top of redemption.

**Verdict on the hypothesis.** Directionally supported in this cohort: both-signals companies retain best (65.96%), beating each single-signal group and neither. Notably it's not "both signals add independently" — the interaction is superadditive relative to the neither baseline (+25.96 vs +6.94 and −8.97).

**What this does and does not prove.**
- Does: in this 2023 signup cohort, first-month breadth-of-giving plus at least one redemption co-occurs with materially higher 24-month retention, and the givers threshold is the load-bearing single signal.
- Does not prove causation. This is observational: activated companies may simply be healthier companies (reverse causality), and confounders like company size or tier aren't controlled — in this extract, nearly all non-retained companies are tier_two and retained ones tier_three, so tier tracks status almost perfectly and is entangled with the signals.
- Also not proven: that 5+ givers is the right threshold (boundary sensitivity untested), that redemption *hurts* alone (small n=29), or that the pattern generalizes beyond 2023-01–2023-07 signups. No confidence intervals or significance tests were run on this data.
arr-rollforward-reconciliationPASS 1.00
Plain-English question: Does CRM ARR match billing ARR — and if not, exactly where is the variance?
Prompt — what the model saw
Reconcile ARR between CRM company records and billing subscriptions as of 2026-09-05. Billing ARR is MRR x 12. Report both totals, the variance, and a decomposition into buckets (status mismatch, rounding, missing records, other) that sums exactly to the variance. List each mismatched account with a suggested owner. Business rule: any subscription with a term other than 12 months must have cf_agreement_end_date populated; list violations.
cost $0.0290 · 118s · in 4,510 / out 18,993 tok
✓ billing_total✓ crm_total✓ variance✓ buckets_sum✓ violations_listed

✓ Correct Answer (ground truth)

Plain-English answer: Billing $604,739.28 vs CRM $603,581.76: a $1,157.52 variance decomposed into status-mismatch, rounding, and missing-record buckets.
The answer in one line: n_companies = 40; n_mismatched_accounts = 7; as_of = 2026-09-05 (+4 more below)
n_companies40
n_mismatched_accounts7
as_of2026-09-05
billing_arr_total604739.28
bucket_sum_equals_variance1
buckets{"status_mismatch": {"amount": "-13158.48", "accounts": [{"company_alias": "C-0C8323BF", "billing_arr": "0.00", "hubspot_arr": "4905.24", "diff": "-4905.24", "suggested_owner": "Dana Mercer"}, {"compa
crm_arr_total603581.76
ruleterm_months != 12 requires cf_agreement_end_date
variance1157.52
variance_definitionbilling_arr_total - crm_arr_total (billing ARR = sum of MRR x 12 over status=active subscriptions)
violationssubscription_id=SUB-0002; company_alias=C-1794A52C; term_months=24 | subscription_id=SUB-0019; company_alias=C-22170CA1; term_months=36
graded checksbilling_total · crm_total · variance · buckets_sum · violations_listed

Model's actual answer

show the model's full answer
# ARR Reconciliation — Billing vs CRM, as of 2026-09-05

**Scope note:** Billing ARR is computed from the 37 subscriptions with status = "active" (2 of the 39 are cancelled and count $0 toward ARR). Both files contain 39 company aliases; 38 match 1:1.

## Totals

- **Billing ARR** = sum(active MRR) × 12 = 50,394.94 × 12 = **$604,739.28**
- **CRM ARR** = sum(hubspot_arr) = **$603,581.76**
- **Variance (billing − CRM) = +$1,157.52**

## Variance decomposition (sums exactly to $1,157.52)

| Bucket | Amount | Detail |
|---|---:|---|
| Status mismatch | −$13,158.48 | 2 cancelled-in-billing but active-ARR in CRM |
| Missing records | +$11,952.00 | 1 billing-only + 1 CRM-only |
| Rounding | −$36.00 | 2 accounts where CRM = billing ARR rounded to nearest $100 |
| Other | +$2,400.00 | 1 account with a real value discrepancy |
| **Total** | **+$1,157.52** | −13,158.48 + 11,952.00 − 36.00 + 2,400.00 |

## Mismatched accounts and suggested owners

No owner/rep field exists in either file, so named owners cannot be assigned from the data; suggestions are by functional owner.

**Status mismatch (CRM should be $0 but isn't):**
- C-0C8323BF — SUB-000E cancelled, billing ARR $0.00 vs CRM $4,905.24 → delta −$4,905.24. Owner: Billing Ops (confirm cancellation propagated to CRM).
- C-0DC4FB8C — SUB-000F cancelled, billing ARR $0.00 vs CRM $8,253.24 → delta −$8,253.24. Owner: Billing Ops.

**Missing records:**
- C-21629AA4 — SUB-0004 active, MRR 2,370.77 × 12 = $28,449.24, no CRM company record → delta +$28,449.24. Owner: Sales/CRM Ops (create CRM record, link subscription).
- C-0D5BBE3A — CRM ARR $16,497.24, no Chargebee subscription at all → delta −$16,497.24. Owner: Billing Ops + Sales/CRM Ops (determine whether the subscription was never created or the CRM record is stale).

**Rounding (CRM ≈ billing rounded to nearest $100):**
- C-0D66DF9E — SUB-0005: 1,932.00 × 12 = $23,184.00 vs CRM $23,200.00 → delta −$16.00. Owner: RevOps (store exact ARR, not rounded).
- C-14D70CE0 — SUB-0008: 1,515.00 × 12 = $18,180.00 vs CRM $18,200.00 → delta −$20.00. Owner: RevOps.

**Other:**
- C-0F7269D7 — SUB-0006: MRR 2,233.00 × 12 = $26,796.00 vs CRM $24,396.00 → delta +$2,400.00. Not a rounding artifact (CRM value isn't a $100 multiple); CRM looks like it was computed from MRR ≈ 2,033.00. Owner: Billing Ops to confirm which MRR is correct.

## Term rule violations (term ≠ 12 months requires cf_agreement_end_date)

4 subscriptions have a non-12-month term; 2 violate:

- **SUB-0002** (C-1794A52C) — 24-month term, cf_agreement_end_date blank — **VIOLATION**
- **SUB-000C** (C-0DB48281) — 24-month term, cf = 2027-11-30 — compliant
- **SUB-0019** (C-22170CA1) — 36-month term, cf_agreement_end_date blank — **VIOLATION**
- **SUB-001A** (C-0FC4DBB8) — 36-month term, cf = 2027-11-30 — compliant

**Methodology notes:** Cancelled subscriptions (SUB-000E, SUB-000F) were excluded from billing ARR; if you instead count all 39 subscriptions, billing ARR would be $617,897.76 and the reconciliation would change materially. No timestamp in the files confirms a 2026-09-05 snapshot — that as-of date is taken as given from your request, not from the data.
kvm-month-over-monthPASS 1.00
Plain-English question: How did the four core KVMs move in August versus July?
Prompt — what the model saw
Report the four core KVMs for 2026-08 against 2026-07: giving rate, redemptions per user, 1:1 meetings engagement, pulse check engagement. For each: value, prior value, absolute and relative change, direction. Name the KVM with the largest relative move and the segment (plan_tier or size_band) driving it if the data supports that; otherwise say the data does not. Under 200 words plus one table.
cost $0.0172 · 52s · in 6,193 / out 8,971 tok
✓ four_kvms✓ biggest_mover_named✓ direction_correct✓ segment_named✓ word_cap

✓ Correct Answer (ground truth)

Plain-English answer: Giving rate 60.27% (up 0.07%), redemptions per user 1.7302 (up 0.01%), with 1:1 meetings and pulse engagement each up fractionally — direction reported for all four.
The answer in one line: n_company_rows_per_month = 30; aggregation = simple mean of company-level values per month
n_company_rows_per_month30
aggregationsimple mean of company-level values per month
biggest_moverpulse_engagement
driving_segmentsize_band=enterprise; 2026-07=0.55; 2026-08=0.2743; rel_change_pct=-50.13
kvms{"giving_rate": {"2026-07": 0.6023, "2026-08": 0.6027, "abs_change": 0.0004, "rel_change_pct": 0.07, "direction": "up"}, "redemptions_per_user": {"2026-07": 1.73, "2026-08": 1.7302, "abs_change": 0.00
segment_detail{"smb": {"2026-07": 0.6588, "2026-08": 0.6573, "rel_change_pct": -0.23}, "mid_market": {"2026-07": 0.593, "2026-08": 0.5942, "rel_change_pct": 0.2}, "enterprise": {"2026-07": 0.55, "2026-08": 0.2743,
graded checksfour_kvms · biggest_mover_named · direction_correct · segment_named · word_cap

Model's actual answer

show the model's full answer
Four core KVMs, 2026-08 vs 2026-07 (simple average across 30 companies):

| KVM | Aug | Jul | Abs Δ | Rel Δ | Direction |
|---|---|---|---|---|---|
| Giving rate | 0.6027 | 0.6023 | +0.0004 | +0.07% | ↑ flat |
| Redemptions per user | 1.7302 | 1.7300 | +0.0002 | +0.01% | ↑ flat |
| 1:1 engagement | 0.4472 | 0.4469 | +0.0003 | +0.07% | ↑ flat |
| Pulse engagement | 0.5086 | 0.6006 | −0.0920 | −15.31% | ↓ down |

Largest relative move: pulse engagement (−15.31%). Segment driving it: size_band = enterprise. Enterprise pulse fell 0.5500 → 0.2743 (−0.2757, −50.13%); every one of the 10 enterprise companies (C-0B2895EF … C-0BA71F12) dropped from ~0.53–0.57 to ~0.26–0.29, while smb (−0.23%) and mid_market (+0.21%) were stable — so the aggregate decline is fully attributable to enterprise. plan_tier is tier_three for all companies, so it cannot differentiate; the driver is size_band.

Method: unweighted means per month; relative change = (Aug − Jul) ÷ Jul. No other data provided.
rewards-ytd-section0.67
Plain-English question: Produce the redemption section of the weekly rewards report through August.
Prompt — what the model saw
Produce the redemption section of the weekly rewards report through the last completed month (2026-08): redemption count, spend, unique redeemers, redemptions per redeemer, provider mix as percent of spend, top 5 countries by redemptions. Say which month is the last completed month. Provider shares must sum to 100.
cost $0.0582 · 142s · in 43,522 / out 25,358 tok
✓ count✗ spend✗ unique_redeemers✓ month_named✓ provider_mix✓ top_country

✓ Correct Answer (ground truth)

Plain-English answer: 378 redemptions, $27,846 spend, 235 unique redeemers at 1.61 each, custom rewards at 39% of spend, US leading all countries.
The answer in one line: redemption_count = 378; unique_redeemers = 235; redemptions_per_redeemer = 1.61; total_spend_usd = 27846.00
redemption_count378
redemptions_per_redeemer1.61
unique_redeemers235
provider_mix_pct_of_spendTangoCard=11.6; Tremendous=30.54; custom=39.05; Snappy=18.81
top5_countries_by_countcountry=US; count=244 | country=CA; count=24 | country=AU; count=21 | country=GB; count=17 | country=NL; count=17
total_spend_usd27846.00
graded checkscount · spend · unique_redeemers · month_named · provider_mix · top_country

Model's actual answer

show the model's full answer
**Last completed month: August 2026** (data runs through 2026-08-31; September 2026 is partial and excluded).

Basis: redemptions_ytd.csv, 378 rows YTD, all 2026-01-01 → 2026-08-31, no missing fields, 378 distinct redemption_ids (no dupes). August subset: n = 48 (monthly counts 45+36+50+52+43+46+58+48 = 378 ✓). Providers found in data (discovered, no whitelist): TangoCard, Tremendous, custom, Snappy — cited exactly as given.

## Redemption section — August 2026

- Redemption count: 48
- Spend: $3,536.00
- Unique redeemers: 46 (distinct user_key; 2 users redeemed twice, e.g. U-4A0FE6, U-0D807F — deduped, not summed)
- Redemptions per redeemer: 48 ÷ 46 = 1.04

Provider mix (% of spend):

```
Provider      Spend ($)   Redemptions   % of spend (raw)     Rounded
Tremendous     1,145.00       22         1145/3536 = 32.3812%   32%
custom         1,057.00        4         1057/3536 = 29.8925%   30%
Snappy           834.00       11          834/3536 = 23.5860%   24%
TangoCard        500.00       11          500/3536 = 14.1403%   14%
Total          3,536.00       48                               100%
```

Shares sum to 100 exactly (1145+1057+834+500 = 3536 ✓; largest-remainder rounding — custom and Snappy rounded up). "custom" is the provider label as it appears in the file.

Top 5 countries by redemptions:

```
1. US   31
2. CA    7
3. GB    2   \
4. SG    2    \  four-way tie at 2; ranks 3-6 are
5. AU    2    /  arbitrary within the tie
   NL    2   /
```
Tie note: GB, SG, AU, NL each have 2 redemptions — the 5th slot is a coin flip among them; CH (1) and IN (1) follow. Country counts sum: 31+7+2+2+2+2+1+1 = 48 ✓.

Data caveats: counts/redeemers/countries are as-recorded in the file; no company-level filter was applied (all aliases included); spend is amount_usd summed as-is.
customer-success 5 tests
churn-save-eligibilityPASS 1.00
Plain-English question: Which at-risk accounts actually qualify for a churn-save offer, and which play fits each?
Prompt — what the model saw
Which at-risk accounts qualify for a churn-save offer under the documented eligibility rules, what amount is at stake per account and in total, and which play fits each (usage revival, executive touch, commercial concession)? Cite the signal that justifies each play. List accounts that look at risk but do not qualify and why.
cost $0.0100 · 45s · in 1,329 / out 6,223 tok
✓ eligible_set✓ total_at_stake✓ plays_cited✓ noneligible_named✓ rules_applied

✓ Correct Answer (ground truth)

Plain-English answer: Only accounts meeting all three rules qualify — e.g. C-0F6C0F34 at $49,707 (executive touch) and C-0B827671 at $25,365 (usage revival) — with amount and play per account.
The answer in one line: n_accounts = 30; snapshot_date = 2026-09-05
n_accounts30
eligibleaccount_alias=C-0F6C0F34; amount_at_stake=49707.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0B827671; amount_at_stake=25365.00; play=usage revival; justifying_signal=usage_trend_3m=declining over the last 3 months | account_alias=C-0B360C78; amount_at_stake=35748.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 75% - risk is commercial, not adoption | account_alias=C-0B0F1BAB; amount_at_stake=5494.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0CA21961; amount_at_stake=16829.00; play=usage revival; justifying_signal=seat utilization 26% is below 50% | account_alias=C-0E9C27D1; amount_at_stake=41235.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 85% - risk is commercial, not adoption … (+2 more)
non_eligible_at_riskaccount_alias=C-0BC71BDD; health_score=55 | account_alias=C-0BA71F12; health_score=52 | account_alias=C-0F6694C3; health_score=43 | account_alias=C-0BE96399; health_score=54 | account_alias=C-0F876796; health_score=47 | account_alias=C-0FCCD2DF; health_score=43 … (+1 more)
ruleshealth_score < 60, churn_save_eligible_amount > 0, renewal within 120 days of snapshot
snapshot_date2026-09-05
total_amount_at_stake224601.00
graded checkseligible_set · total_at_stake · plays_cited · noneligible_named · rules_applied

Model's actual answer

show the model's full answer
All figures below come only from churnzero_accounts.csv. Eligibility requires all three rules (R1 AND R2 AND R3).

**R3 window arithmetic (snapshot 2026-09-05 + 120 days):**
25 days left in Sep → Oct (31) → Nov (30) → Dec (31) = 25+31+30+31 = 117 days reaches 2026-12-31; +3 = 2027-01-03. Cutoff: renewal_date ≤ 2027-01-03.

---

## 1. Qualifying accounts (8) — total at stake $224,601.00

| Account | Health | Eligible $ | Renewal (days out) | Usage 3m | Seats used/owned (util %) | Champion |
|---|---|---|---|---|---|---|
| C-0B0F1BAB | 38 | 5,494.00 | 2026-09-23 (18) | flat | 238/363 (65.6%) | false |
| C-0E9C27D1 | 39 | 41,235.00 | 2026-09-24 (19) | flat | 134/157 (85.4%) | true |
| C-0F6C0F34 | 51 | 49,707.00 | 2026-10-03 (28) | growing | 308/395 (78.0%) | false |
| C-0B360C78 | 57 | 35,748.00 | 2026-10-28 (53) | growing | 246/327 (75.2%) | true |
| C-0D3278C7 | 54 | 17,602.00 | 2026-11-12 (68) | declining | 126/380 (33.2%) | true |
| C-0B827671 | 56 | 25,365.00 | 2026-11-14 (70) | declining | 113/202 (55.9%) | true |
| C-0CEF69FD | 53 | 32,621.00 | 2026-11-21 (77) | growing | 97/136 (71.3%) | false |
| C-0CA21961 | 58 | 16,829.00 | 2026-12-28 (114) | flat | 84/325 (25.8%) | true |

Total: 49,707 + 25,365 + 35,748 + 5,494 + 16,829 + 41,235 + 32,621 + 17,602 = **$224,601.00**

## 2. Play assignment

The rules file defines eligibility only, not play mapping. I applied this signal rule to the fields provided: declining usage or utilization <40% → usage revival; champion_active=false with no usage decline → executive touch; at-risk health despite healthy adoption (growing usage and/or ≥70% utilization) and an active champion → commercial concession.

**Usage revival — $59,796.00**
- C-0B827671 ($25,365): usage_trend_3m = declining
- C-0D3278C7 ($17,602): usage_trend_3m = declining; utilization 33.2%
- C-0CA21961 ($16,829): utilization 84/325 = 25.8% (worst in set) — flat trend masks heavy shelfware

Subtotal: 25,365 + 17,602 + 16,829 = $59,796.00

**Executive touch — $87,822.00** (all three have champion_active = false)
- C-0F6C0F34 ($49,707): no champion; usage growing, 78.0% utilization — relationship gap, not adoption gap
- C-0CEF69FD ($32,621): no champion; usage growing, 71.3% utilization
- C-0B0F1BAB ($5,494): no champion; flat usage, 65.6% utilization; renews in 18 days — most urgent exec outreach

Subtotal: 49,707 + 32,621 + 5,494 = $87,822.00

**Commercial concession — $76,983.00** (healthy adoption, yet health <60 → value/price lever)
- C-0E9C27D1 ($41,235): health 39 but 85.4% utilization and active champion; renews in 19 days
- C-0B360C78 ($35,748): health 57 but growing usage, 75.2% utilization, active champion

Subtotal: 41,235 + 35,748 = $76,983.00

Cross-check: 59,796 + 87,822 + 76,983 = $224,601.00 ✓

## 3. At-risk (health <60) but NOT qualifying — 7 accounts

Fails R2 (eligible amount = $0.00):
- C-0BC71BDD (health 55, renews 2026-10-27 — would pass R3)
- C-0BE96399 (health 54, renews 2026-10-29 — would pass R3)
- C-10A56B0F (health 54, renews 2026-12-12 — would pass R3)
- C-0F6694C3 (health 43, renews 2027-03-21 — also fails R3)
- C-0FCCD2DF (health 43, renews 2027-04-23 — also fails R3)

Fails R3 only (renewal beyond 2027-01-03):
- C-0F876796 (health 47, eligible $19,958.00, renews 2027-02-06 = 154 days out: 25+31+30+31+31+6)
- C-0BA71F12 (health 52, eligible $6,824.00, renews 2027-04-11 = 218 days out: 25+31+30+31+31+30+31+9)

Caveat: the concession vs. usage vs. exec mapping is my analyst judgment on the five provided fields; the data contains no documented play-definition table.
expansion-kitPASS 1.00
Plain-English question: Build the expansion kit for the customer who just replied to our outreach.
Prompt — what the model saw
A customer replied to outreach (reply included). Build the expansion kit: seat coverage (licensed over headcount), two lines on usage health, headroom in seats and ARR at the current per-seat rate, who replied and whether they can buy, the right buyer from our contacts, and a reply email under 150 words that moves toward the right person without being pushy and cites one real usage fact.
cost $0.0052 · 28s · in 429 / out 2,675 tok
✓ coverage_pct✓ headroom_arr✓ right_buyer✓ replier_not_buyer✓ usage_fact✓ email_under_150

✓ Correct Answer (ground truth)

Plain-English answer: 37.5% seat coverage, 250 headroom seats worth $15,000 ARR at $60 a seat, usage up 43.2% in six months — and the replier cannot buy; Dana R. (VP People) is the right buyer.
The answer in one line: headroom_seats = 250; coverage_pct = 37.5; company_alias = C-0DDFC9A7 (+3 more below)
headroom_seats250
coverage_pct37.5
company_aliasC-0DDFC9A7
headroom_arr15000.00
per_seat_annual_rate60.00
replierMaria S.
replier_can_buy0
right_buyername=Dana R.; title=VP People
usage_factsactive_users_2026_08=126; active_users_2026_03=88; growth_pct_6m=43.2; licensed_seat_utilization_pct=84
usage_health_linesMonthly active users grew from 88 (2026-03) to 126 (2026-08), +43.2% over 6 months., Last month 126 of 150 licensed seats were active (84.0% utilization), approaching the seat cap.
graded checkscoverage_pct · headroom_arr · right_buyer · replier_not_buyer · usage_fact · email_under_150

Model's actual answer

show the model's full answer
# Expansion Kit — C-0DDFC9A7

## 1. Seat coverage
Licensed 150 ÷ headcount 400 = **37.5% of the company is licensed**. 250 employees (400 − 150) have no seat.

## 2. Usage health
- MAU climbed every month without exception: 88 (Mar) → 126 (Aug), +38 users = +43.2% (38 ÷ 88) in 5 months, at a steady ~7–8 net adds/month (7, 7, 8, 8, 8).
- Utilization of the license is 126 ÷ 150 = **84%**; at the recent +8/month pace, MAU reaches the 150-seat cap around Nov 2026 (Sep 134, Oct 142, Nov 150).

## 3. Headroom
- **Seats:** 150 licensed vs 400 headcount = 250 unlicensed seats.
- **Per-seat rate:** $9,000 ARR ÷ 150 seats = **$60/seat/year**.
- **ARR headroom:** 250 × $60 = **$15,000 incremental ARR** (potential $24,000 total, 2.7× current). Near-term trigger: only 24 seats of capacity above current MAU.
- Note: MAU is a usage proxy, not a seat-assignment count — data for licensed-seat-active users is not provided.

## 4. Who replied, and can they buy?
**Maria S., People Operations Coordinator** (last engaged 2026-09-02, same day as the reply). She cannot buy — she explicitly defers purchasing. She is a strong champion: confirms adoption in her own words and offers an introduction, noting Dana has been asking about usage numbers.

## 5. Right buyer
**Dana R., VP People** — named by Maria as owner of budget and seat expansion. Caveat: her last engagement was 2026-05-18 (~3.5 months before the reply), so the relationship needs a re-entry point; Maria's offer is the natural bridge.

## 6. Reply email (96 words)

> Subject: RE: Growing your team's recognition program
>
> Hi Maria,
>
> Great to hear the team is loving Bonusly — thanks for the note.
>
> One thing that jumped out on our side: monthly active users are up from 88 in March to 126 in August (+43%), which puts you at 84% of your 150-seat license.
>
> Since budget and seats sit with Dana, whatever is easiest for you: I can send a short usage summary you could forward to her, or happily take you up on that intro.
>
> Thanks!
> Cole
onboarding-statusPASS 1.00
Plain-English question: What is done and what is missing in this customer's onboarding before the mid-onboarding call?
Prompt — what the model saw
Prep the CSM for the mid-onboarding call: what is complete (integrations connected, allowance set, admins added, first recognition given, first redemption), what is not, early engagement signals, and the three things to cover on the call. Mark nothing complete without a data field that shows it.
cost $0.0048 · 28s · in 439 / out 2,320 tok
✓ complete_items✓ incomplete_items✓ three_items✓ grounded

✓ Correct Answer (ground truth)

Plain-English answer: 25 days live: Slack connected, allowance set, admins added — but the HRIS integration is not connected, with early engagement signals and three call priorities.
The answer in one line: days_live = 25; company_alias = C-0D284E42
days_live25
call_itemsConnect the HRIS integration (integration_hris is empty) to automate user provisioning and offboarding., Get to first redemption (first_redemption_at is empty): walk admins through the rewards catalog and confirm allowance funding., Build the giving habit: daily active givers averaged 4.3 in week 1 vs 13.0 last week - agree on manager nudges and a recognition cadence.
checklist{"slack_integration_connected": {"complete": true, "field": "integration_slack", "value": "2026-08-12"}, "hris_integration_connected": {"complete": false, "field": "integration_hris", "value": ""}, "a
company_aliasC-0D284E42
engagement_signalsavg_active_givers_week1=4.3; avg_active_givers_last_week=13; latest_active_givers=15; days_with_activity=25
graded checkscomplete_items · incomplete_items · three_items · grounded

Model's actual answer

show the model's full answer
CSM PREP — C-0D284E42 (signup 2026-08-11; data through 2026-09-04)

WHAT IS COMPLETE (each backed by a data field)
- Slack integration: connected 2026-08-12 (integration_slack) — 1 day after signup.
- Allowance set: 2026-08-13 (allowance_set) — 2 days after signup.
- Admins added: 2 (admins_added = 2).
- First recognition given: 2026-08-15 14:22 (first_recognition_at) — 4 days after signup.

WHAT IS NOT COMPLETE
- HRIS integration: integration_hris is empty — no connection date on record.
- First redemption: first_redemption_at is empty — no redemption recorded through the last data point (2026-09-04), i.e., 24 days post-signup (Aug 11 → Sep 4: 20 + 4 = 24 days).

EARLY ENGAGEMENT SIGNALS (onboarding_usage, 2026-08-11 → 2026-09-04, 25 days)
- Daily active givers grew 3 → 15 (15 / 3 = 5.0x over the window).
- Weekly averages: W1 (08-11–08-17): 30/7 = 4.3; W2 (08-18–08-24): 53/7 = 7.6; W3 (08-25–08-31): 75/7 = 10.7; W4 partial (09-01–09-04): 56/4 = 14.0. Week-over-week: 53/30 = +77%, 75/53 = +42%, 14.0/10.7 = +31% (W4 is a 4-day partial; compared on daily average only).
- No zero-usage days; minimum was 3 active givers (08-11, 08-12).
- Data limitation: usage file counts active givers only. It contains nothing on recipients, redemption activity, or point balances, so the redemption gap cannot be diagnosed from this data — only observed via the empty first_redemption_at field.

THREE THINGS TO COVER ON THE CALL
1. HRIS connection (not complete): walk through connecting it and what it unlocks (automated user provisioning/sync). Note admins are currently manual — only 2 admins added.
2. First redemption (not complete): it's the last onboarding milestone untouched. Demo the rewards catalog and confirm users know how to redeem and that point balances/allowances are landing as expected. Flag that we have no visibility into why it hasn't happened (no redemption/balance fields in the data).
3. Sustain and broaden momentum: giver activity is up 5x in 25 days with no dead days — recognize it, then set a concrete next milestone (e.g., first redemption date and a giver-count target) before hanging up. Given admins_added = 2, also ask whether more admins are needed as usage scales.
renewal-risk-conflicting-datesPASS 1.00
Plain-English question: Two systems disagree on renewal dates — which do you trust per account, and who is actually renewing inside 90 days?
Prompt — what the model saw
Produce a 90-day renewal risk brief. Two systems hold renewal dates and disagree on some accounts; decide which to trust per account and say why (multi-year contracts are known to be wrong in ChurnZero). For every renewal: company, CSM, ARR, date used, seat utilization, 3-month usage trend, risk rating with one sentence of evidence. Flag every disagreement. Close with total ARR renewing and ARR at risk.
cost $0.0259 · 113s · in 5,065 / out 17,849 tok
✓ total_renewing✓ arr_at_risk✓ disagreements_flagged✓ trust_rule

✓ Correct Answer (ground truth)

Plain-English answer: Chargebee wins per account (multi-year contracts are known wrong in ChurnZero), with each trusted date, the reason, and a 90-day-window flag.
The answer in one line: n_accounts = 20; n_disagreements = 5; snapshot_date = 2026-09-05 (+2 more below)
n_accounts20
n_disagreements5
accounts20 items — e.g. account_alias=C-0B144C78; csm=Cole Ingram; arr=30899.00; trusted_renewal_date=2026-11-02; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.4; usage_3m_ratio=1.03; risk=low; evidence=3-month usage ratio 1.03 (last3 avg 103 vs prior3 100), seat utilization 75% | account_alias=C-0B20DB64; csm=Dana Mercer; arr=21770.00; trusted_renewal_date=2026-10-07; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=56.6; usage_3m_ratio=1; risk=medium; evidence=3-month usage ratio 1.00 (last3 avg 295 vs prior3 295), seat utilization 57% | account_alias=C-0B344485; csm=Elena Sinclair; arr=64384.00; trusted_renewal_date=2026-11-16; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=78; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 241 vs prior3 231), seat utilization 78% | account_alias=C-0B369871; csm=Dana Mercer; arr=85128.00; trusted_renewal_date=2026-10-29; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.1; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 330 vs prior3 316), seat utilization 75%
arr_at_risk_90d359409.00
disagreementsaccount_alias=C-0B7D2C30; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-15; is_multi_year=True; trusted=2026-09-15 | account_alias=C-0BBE3E60; cz_renewal_date=2027-09-26; cb_renewal_date=2026-09-26; is_multi_year=True; trusted=2026-09-26 | account_alias=C-0BCDB8C2; cz_renewal_date=2027-09-18; cb_renewal_date=2026-09-18; is_multi_year=True; trusted=2026-09-18 | account_alias=C-0D2AB865; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-22; is_multi_year=True; trusted=2026-09-22 | account_alias=C-0F5D2323; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-29; is_multi_year=True; trusted=2026-09-29
snapshot_date2026-09-05
total_arr_renewing_90d1048715.00
trust_rulemulti-year contracts: Chargebee is authoritative (ChurnZero known wrong); otherwise systems agree or Chargebee wins
window2026-09-05 to 2026-12-04
graded checkstotal_renewing · arr_at_risk · disagreements_flagged · trust_rule

Model's actual answer

show the model's full answer
90-DAY RENEWAL RISK BRIEF
=========================

METHOD & ASSUMPTIONS (stated up front)
- As-of date: 2026-09-01 (latest usage month in data is 2026-08). 90-day window = 2026-09-01 to 2026-11-30. All 20 renewals (as dated below, 2026-09-15 through 2026-11-24) fall inside it; none excluded.
- Date arbitration rule: where ChurnZero (CZ) and Chargebee (CB) disagree, the account is multi-year in CB, and multi-year dates are known-wrong in CZ — therefore CB is trusted for all 5 conflicts. The 15 annual accounts have identical dates in both systems (verified row by row).
- Seat utilization = seats_used / seats (ChurnZero snapshot). 3-month usage trend = active_users 2026-06 to 2026-08 (latest 3 months available).
- Risk rules: HIGH = 3-month active-user decline >=10% OR utilization <30%. MEDIUM = utilization <60% with usage flat (within ±5%). LOW = utilization >=60% and usage change <10% either way.
- Data completeness: all 20 accounts appear in all three files with full 12-month usage; no missing fields. Note: "seats_used" (CZ) and "active_users" (usage file) are different measures and do not reconcile (e.g., C-0B20DB64: 294 active vs 214 seats used); each is used only for its defined purpose above.

DATE CONFLICTS — ALL 5 FLAGGED (all are the multi-year accounts; zero annual-term conflicts)
- C-0B7D2C30: CZ 2026-09-10 vs CB 2026-09-15 (36-mo, multi-year) -> use CB 2026-09-15
- C-0BCDB8C2: CZ 2027-09-18 vs CB 2026-09-18 (36-mo, multi-year) -> use CB 2026-09-18 (CZ date would push this $54,427 out of the 90-day window entirely)
- C-0D2AB865: CZ 2026-09-10 vs CB 2026-09-22 (24-mo, multi-year) -> use CB 2026-09-22
- C-0BBE3E60: CZ 2027-09-26 vs CB 2026-09-26 (24-mo, multi-year) -> use CB 2026-09-26 (CZ date would remove this $30,993 from the window)
- C-0F5D2323: CZ 2026-09-10 vs CB 2026-09-29 (24-mo, multi-year) -> use CB 2026-09-29
Pattern: CZ errs both directions on multi-year (5-19 days early on three accounts; a full year+ late on two), consistent with the known defect.

RENEWALS (chronological by date used)

 1. C-0B7D2C30 | Dana Mercer | $65,901 | 2026-09-15 (CB; CONFLICT: CZ 09-10) | util 57.6% (274/476) | Jun-Aug trend -13.4% (97->84) | HIGH — active users down 13.4% in 3 months on 57.6% seat utilization.
 2. C-0BCDB8C2 | Cole Ingram | $54,427 | 2026-09-18 (CB; CONFLICT: CZ 2027-09-18) | util 54.7% (232/424) | trend -13.4% (127->110) | HIGH — 13.4% active-user drop in 3 months on 54.7% utilization.
 3. C-0D2AB865 | Elena Sinclair | $38,022 | 2026-09-22 (CB; CONFLICT: CZ 09-10) | util 61.4% (250/407) | trend -12.8% (125->109) | HIGH — 12.8% active-user drop in 3 months on 61.4% utilization.
 4. C-0BBE3E60 | Dana Mercer | $30,993 | 2026-09-26 (CB; CONFLICT: CZ 2027-09-26) | util 64.9% (74/114) | trend -15.4% (39->33) | HIGH — steepest decline in the book, -15.4% in 3 months, on 64.9% utilization.
 5. C-0F5D2323 | Cole Ingram | $90,647 | 2026-09-29 (CB; CONFLICT: CZ 09-10) | util 28.5% (111/390) | trend -15.0% (20->18) | HIGH — only 279 of 390 paid seats in use (28.5%) plus a 15.0% usage drop on the largest multi-year ARR.
 6. C-0EC6999D | Elena Sinclair | $79,419 | 2026-10-03 (CZ=CB) | util 27.7% (31/112) | trend -11.8% (17->15) | HIGH — 72% of paid seats idle (31/112) and usage down 11.8%.
 7. C-0B20DB64 | Dana Mercer | $21,770 | 2026-10-07 (CZ=CB) | util 56.6% (214/378) | trend 0.0% (294->294) | MEDIUM — 164 idle seats (56.6% utilization) with flat usage leaves shrink room at renewal.
 8. C-0BBC4E7A | Cole Ingram | $56,374 | 2026-10-10 (CZ=CB) | util 67.7% (228/337) | trend -2.1% (142->139) | LOW — stable usage on two-thirds utilization.
 9. C-0FD551AB | Elena Sinclair | $48,815 | 2026-10-14 (CZ=CB) | util 55.9% (210/376) | trend +2.4% (123->126) | MEDIUM — 166 idle seats (55.9% utilization), though usage is stable-to-growing.
10. C-0F9F8F13 | Dana Mercer | $46,230 | 2026-10-18 (CZ=CB) | util 56.5% (199/352) | trend -1.6% (185->182) | MEDIUM — 153 idle seats (56.5% utilization) with flat usage.
11. C-0BC34584 | Cole Ingram | $16,740 | 2026-10-22 (CZ=CB) | util 66.2% (327/494) | trend +1.9% (104->106) | LOW — stable usage on 66.2% utilization.
12. C-0B7A7546 | Elena Sinclair | $35,062 | 2026-10-25 (CZ=CB) | util 88.8% (182/205) | trend -1.6% (64->63) | LOW — highest utilization in the book with flat usage.
13. C-0B369871 | Dana Mercer | $85,128 | 2026-10-29 (CZ=CB) | util 75.1% (317/422) | trend +2.2% (326->333) | LOW — usage growing on 75.1% utilization.
14. C-0B144C78 | Cole Ingram | $30,899 | 2026-11-02 (CZ=CB) | util 75.4% (169/224) | trend +5.0% (101->106) | LOW — usage growing on 75.4% utilization.
15. C-0FC4DBB8 | Elena Sinclair | $94,732 | 2026-11-05 (CZ=CB) | util 76.7% (356/464) | trend +2.1% (189->193) | LOW — usage growing on 76.7% utilization (largest ARR in book, healthy).
16. C-0D5BBE3A | Dana Mercer | $39,740 | 2026-11-09 (CZ=CB) | util 83.3% (85/102) | trend +3.4% (88->91) | LOW — usage growing on 83.3% utilization.
17. C-0FB9D5AF | Cole Ingram | $63,158 | 2026-11-13 (CZ=CB) | util 72.4% (144/199) | trend +1.7% (173->176) | LOW — usage growing on 72.4% utilization.
18. C-0B344485 | Elena Sinclair | $64,384 | 2026-11-16 (CZ=CB) | util 78.0% (224/287) | trend +2.5% (238->244) | LOW — usage growing on 78.0% utilization.
19. C-0CB2C1B4 | Dana Mercer | $40,628 | 2026-11-20 (CZ=CB) | util 81.6% (386/473) | trend +4.3% (47->49) | LOW — usage growing on 81.6% utilization.
20. C-22170CA1 | Cole Ingram | $45,646 | 2026-11-24 (CZ=CB) | util 85.4% (251/294) | trend +2.1% (143->146) | LOW — usage growing on 85.4% utilization.

TOTALS (arithmetic shown)
- Total renewing ARR (all 20): 65,901 + 54,427 + 38,022 + 30,993 + 90,647 + 79,419 + 21,770 + 56,374 + 48,815 + 46,230 + 16,740 + 35,062 + 85,128 + 30,899 + 94,732 + 39,740 + 63,158 + 64,384 + 40,628 + 45,646 = $1,048,715.00
- ARR at risk (HIGH, 6 accounts): 65,901 + 54,427 + 38,022 + 30,993 + 90,647 + 79,419 = $359,409.00 (34.3% of total; 359,409 / 1,048,715 = 34.27%)
- MEDIUM (3 accounts): 21,770 + 48,815 + 46,230 = $116,815.00 (11.1%)
- LOW (11 accounts): $572,491.00 (54.6%)
- If MEDIUM is included, ARR at risk = 359,409 + 116,815 = $476,224.00 (45.4%)

Note: all 6 HIGH accounts renew in the first 34 days of the window (2026-09-15 to 2026-10-03) — the riskiest $359k has the least runway, and all six are also inside the conflicted-date set or the sub-30%-utilization pair.
ticket-theme-synthesis0.80
Plain-English question: What are the real themes in this quarter's support tickets, ignoring the existing tags?
Prompt — what the model saw
Synthesize the quarter's support tickets into 4 to 7 themes. Do not trust existing tags; read the text. Per theme: name, count, share, distinct accounts, ARR affected, two ticket ids, one-line recommendation. Separate broad patterns from single-account noise and rank by ARR exposure, not volume.
cost $0.0363 · 135s · in 9,722 / out 23,130 tok
✗ theme_count_range✓ counts_verifiable✓ single_account_flagged✓ arr_ranked✓ ticket_ids

✓ Correct Answer (ground truth)

Plain-English answer: 80 tickets in 4-7 themes — points-not-posting leads (20 tickets, 25%, $31,100 ARR affected) — each with counts, distinct accounts, ARR, and a recommendation.
The answer in one line: total_tickets = 80; quarter_window = 2026-06-01 through 2026-08-31
total_tickets80
noteexisting tags are noisy/unreliable; classify by text
quarter_window2026-06-01 through 2026-08-31
ranking_by_arr_exposurehris_provisioning, redemption_failures, billing_dispute, points_not_posting, slack_sync
sample_ticket_ids80 items — e.g. IC-460004, IC-460016, IC-460001, IC-460041
single_account_noise_themebilling_dispute
themes{"points_not_posting": {"count": 20, "share_pct": 25.0, "distinct_accounts": 9, "arr_affected": "31100.00", "example_ticket_ids": ["IC-460001", "IC-460002"], "single_account": false}, "slack_sync": {"
graded checkstheme_count_range · counts_verifiable · single_account_flagged · arr_ranked · ticket_ids

Model's actual answer

show the model's full answer
Q2 support-theme synthesis. All 80 tickets (2026-06-01 → 2026-08-29), 24 distinct accounts, total account ARR $284,800. Tags were ignored (verified unreliable: e.g., IC-460020 is tagged "billing" but is a points-delivery ticket; IC-460050 "billing" is a Slack ticket). "ARR affected" = sum of full ARR of distinct accounts filing under the theme — the data does not allow partial-exposure estimates.

Themes ranked by ARR exposure:

1. HRIS provisioning failures — new hires not provisioned
   Count: 12 (15.0% of tickets) | Distinct accounts: 3 (C-0B2213A9 $36,000, C-0DDFC9A7 $48,000, C-0F6C0F34 $30,000) | ARR affected: $114,000 (40.0% of portfolio ARR: 36,000+48,000+30,000)
   Tickets: IC-460060, IC-460053
   Recommendation: Escalate as P1 to the three largest accounts; fix the silent-failure mode first ("provisioning log shows no errors" while syncs skip new hires) — this is your biggest ARR at risk.

2. Redemption / gift card checkout failures
   Count: 18 (22.5%) | Distinct accounts: 7 (C-14264ABD $11,000, C-0B827671 $10,700, C-0B0F1BAB $10,300, C-0D9CA315 $9,600, C-0FCCD2DF $9,600, C-0CEF69FD $8,900, C-0F876796 $8,700) | ARR affected: $68,800 (24.2%)
   Tickets: IC-460024, IC-460029
   Recommendation: Highest ticket volume; fix checkout timeout and add automatic point reversal when orders error — "points deducted but no gift card" is a trust-destroying failure.

3. Invoice accuracy — seat counts and renewal pricing (SINGLE-ACCOUNT PATTERN)
   Count: 16 (20.0%) | Distinct accounts: 1 (C-0E9C27D1 $52,000) | ARR affected: $52,000 (18.3%)
   Tickets: IC-460069, IC-460078
   Recommendation: Do not treat as a product theme — this is one account escalation: C-0E9C27D1 filed 16 tickets in 11 weeks over being charged 200 seats vs 150 licensed and wrong-tier renewal; audit their invoice, credit the delta, and assign an exec sponsor before renewal.

4. Points balance not posting / stale balance
   Count: 12 (15.0%) | Distinct accounts: 7 (C-0D0B047C $4,500, C-0BF20542 $4,500, C-0D6CC8E3 $4,200, C-0D284E42 $3,400, C-21FEBCBB $2,900, C-0BE96399 $2,700, C-0DD0626C $2,500) | ARR affected: $24,700 (8.7%)
   Tickets: IC-460008, IC-460017
   Recommendation: Cluster of small accounts reporting stale balances after weekends ("since Tuesday") — investigate weekend batch-processing delays in the points ledger.

5. Slack integration breakage
   Count: 14 (17.5%) | Distinct accounts: 4 (C-10A56B0F $5,400, C-8C2E8F00 $5,200, C-0B843542 $4,400, C-0BA71F12 $3,900) | ARR affected: $18,900 (6.6%)
   Tickets: IC-460049, IC-460046
   Recommendation: Fix the OAuth token-refresh bug (re-auth "does not stick," toggle resets) and the team-wide slash-command error, then proactive-notify affected admins.

6. Sent recognitions show delivered but points never arrive
   Count: 8 (10.0%) | Distinct accounts: 5 (C-0D3278C7 $3,500, C-0D284E42 $3,400, C-0B2895EF $2,900, C-0BE96399 $2,700, C-0DD0626C $2,500) | ARR affected: $15,000 (5.3%)
   Tickets: IC-460004, IC-460011
   Recommendation: Investigate the recognition delivery pipeline where senders see "delivered" status with no points posted — the false "delivered" state masks the failure from monitoring.

Notes:
- Broad patterns (multi-account): themes 1, 2, 4, 5, 6. Single-account noise: theme 3 (C-0E9C27D1 only, but 20% of all ticket volume).
- Theme ARR totals ($293,400) exceed the $284,800 portfolio by $8,600 because three small accounts (C-0D284E42, C-0BE96399, C-0DD0626C) filed under both point-ledger themes 4 and 6 — same underlying ledger, different symptoms.
- Volume vs ARR divergence: redemption is #1 by volume (18 tickets) but its accounts are all ≤$11,000 ARR; HRIS is #1 by ARR because it hits the three largest accounts ($30k–$48k). Billing is #2 by volume but is a single account.
marketing 5 tests
case-study-matchPASS 1.00
Plain-English question: Which three existing customers are the best case-study matches for this prospect?
Prompt — what the model saw
Given the prospect, rank the three most similar existing customers with public case studies for use as social proof. Similarity on industry, size band, use case, region. Cite the fields that drove each match. Do not name any customer without a case study and never include billing data or contact names.
cost $0.0057 · 25s · in 710 / out 2,941 tok
✓ top3✓ decoys_excluded✓ cites_fields✓ no_billing

✓ Correct Answer (ground truth)

Plain-English answer: Three ranked matches on industry, size, use case, and region (e.g. C-11C31562), with the higher-scoring decoy explicitly excluded for a mismatched field.
The answer in one line: prospect_alias = C-82AF3719
constraintsno_billing_data, no_contact_names, only_customers_with_has_case_study_true_named
decoys_excludedcompany_alias=C-D6217CAA; similarity_score=4; reason=no_case_study | company_alias=C-C153868F; similarity_score=4; reason=no_case_study
prospect_aliasC-82AF3719
prospect_fields_usedindustry, size_band, use_case, region
top3company_alias=C-11C31562; similarity_score=3 | company_alias=C-64171065; similarity_score=3 | company_alias=C-A13C193D; similarity_score=3
graded checkstop3 · decoys_excluded · cites_fields · no_billing

Model's actual answer

show the model's full answer
Ranking of the three most similar case-study customers for prospect C-82AF3719 (Technology / Mid-Market / employee_recognition / NA-West).

**Method.** Eligible pool = the 8 customers with has_case_study=true. Similarity = count of exact field matches on the four provided dimensions (industry, size_band, use_case, region), max 4/4.

**Scoring (exact-match count vs. prospect):**

| company_alias | industry | size_band | use_case | region | Score |
|---|---|---|---|---|---|
| C-11C31562 | ✗ (Manufacturing) | ✓ Mid-Market | ✓ employee_recognition | ✓ NA-West | 3/4 |
| C-64171065 | ✓ Technology | ✓ Mid-Market | ✓ employee_recognition | ✗ (NA-East) | 3/4 |
| C-A13C193D | ✓ Technology | ✓ Mid-Market | ✗ (retention) | ✓ NA-West | 3/4 |
| C-9D69D147 | ✗ | ✓ Mid-Market | ✗ (retention) | ✓ NA-West | 2/4 |
| C-CD4829A7 | ✓ Technology | ✓ Mid-Market | ✗ (retention) | ✗ (NA-East) | 2/4 |
| C-FEFCC020 | ✗ | ✗ | ✗ (wellness) | ✗ | 0/4 |
| C-11209339 | ✗ | ✗ | ✗ (peer_milestones) | ✗ | 0/4 |
| C-0CE315B2 | ✗ | ✗ | ✗ (wellness) | ✗ | 0/4 |

Three customers tie at 3/4. Tie-break: use_case weighted highest (the buyer's specific problem), then region, then industry (softest signal for a horizontal recognition product). Under that ordering:

**1. C-11C31562** — 3/4
- Matches: size_band=Mid-Market, use_case=employee_recognition, region=NA-West
- Differs: industry=Manufacturing (prospect: Technology)

**2. C-64171065** — 3/4
- Matches: industry=Technology, size_band=Mid-Market, use_case=employee_recognition
- Differs: region=NA-East (prospect: NA-West)

**3. C-A13C193D** — 3/4
- Matches: industry=Technology, size_band=Mid-Market, region=NA-West
- Differs: use_case=retention (prospect: employee_recognition)

**Notes and limitations:**
- All three named customers have has_case_study=true; no customer without a case study is named.
- No billing data or contact names included — none provided in the source, none inferred.
- Similarity is limited to the four fields provided. Data on company headcount, tenure, deployment scale, or outcomes was not provided and is not claimed.
- The 1–3 ordering among the three 3/4 ties rests on the stated weighting; if industry were weighted above region, C-64171065 and C-11C31562 would swap.
channel-attribution-mathPASS 1.00
Plain-English question: Compute cost per SQM, cost per SQO, and pipeline per dollar for every channel over six months.
Prompt — what the model saw
Compute channel performance for the trailing 6 months. Paid: spend, SQMs, SQOs, cost per SQM, cost per SQO, SQM-to-SQO rate, pipeline amount, pipeline per dollar. Organic: volume, SQO rate, pipeline. Report a channel with spend and zero SQMs as undefined, not zero. Flag rows where SQO date precedes SQM date. Finish with a reallocation recommendation and your confidence given sample sizes.
cost $0.0204 · 65s · in 8,611 / out 8,852 tok
✓ undefined_channel✓ undefined_not_zero✓ flagged_rows✓ best_channel✓ reallocation

✓ Correct Answer (ground truth)

Plain-English answer: LinkedIn Ads at $960/SQM and $3,000/SQO, organic search at $90,000 pipeline on zero spend, paid social flagged at $18,000 with zero SQMs, and a channel winner named.
The answer in one line: best_pipeline_per_dollar_channel = paid_search
best_pipeline_per_dollar_channelpaid_search
confidence_notewebinars and referral have small SQM samples (<20); treat rates as directional.
flagged_rows_sqo_before_sqmCT-000041, CT-000044
months2026-03, 2026-04, 2026-05, 2026-06, 2026-07, 2026-08
per_channel{"linkedin_ads": {"spend_usd": 24000, "sqms": 25, "sqos": 8, "pipeline_amount": 96000, "cost_per_sqm": 960.0, "cost_per_sqo": 3000.0, "sqm_to_sqo_rate": 0.32, "pipeline_per_dollar": 4.0}, "organic_sea
reallocation_recommendationfrom=paid_social; to=paid_search; rationale=paid_social has spend but zero SQMs (cost per SQM undefined, not zero); paid_search has the highest pipeline per dollar.
undefined_channelspaid_social
graded checksundefined_channel · undefined_not_zero · flagged_rows · best_channel · reallocation

Model's actual answer

show the model's full answer
# Channel Performance — Trailing 6 Months (Mar–Aug 2026)

All arithmetic below was computed from the two files as given. Coverage: 122 contact rows, spend across 6 months × 4 paid channels. Total paid spend = $87,000.

## Paid channels

| Channel | Spend | SQMs | SQOs | Cost/SQM | Cost/SQO | SQM→SQO | Pipeline | Pipe/$ |
|---|---|---|---|---|---|---|---|---|
| paid_search | $36,000 (6×$6,000) | 40 | 18 | $900.00 | $2,000.00 | 45.0% | $720,000 | $20.00 |
| linkedin_ads | $24,000 (6×$4,000) | 25 | 8 | $960.00 | $3,000.00 | 32.0% | $96,000 | $4.00 |
| paid_social | $18,000 (6×$3,000) | 0 | 0 | **undefined** | **undefined** | **undefined** | $0 | $0.00 |
| webinars | $9,000 (6×$1,500) | 12 | 5 | $750.00 | $1,800.00 | 41.7% | $60,000 | $6.67 |

Arithmetic examples: paid_search cost/SQM = 36,000/40 = $900; SQM→SQO = 18/40 = 45%; pipe/$ = 720,000/36,000 = $20. paid_social has spend and zero first-touch contacts, so per-SQM/per-SQO metrics are undefined, not zero (per your rule). Pipeline/$ for paid_social = $0/$18,000 = $0.00.

## Organic channels (no spend; spend-based metrics not applicable)

| Channel | Volume | SQOs | SQO rate | Pipeline |
|---|---|---|---|---|
| organic_search | 30 | 10 | 33.3% | $90,000 |
| referral | 15 | 6 | 40.0% | $48,000 |

## Data-quality flags — SQO date precedes SQM date

Both inverted rows are linkedin_ads, each carrying $12,000 pipeline:

- CT-000041 (linkedin_ads): SQM 2026-06-14, SQO 2026-06-09 — SQO precedes SQM by 5 days
- CT-000044 (linkedin_ads): SQM 2026-07-23, SQO 2026-07-18 — SQO precedes SQM by 5 days

Sensitivity: excluding these two rows, linkedin_ads drops to 6/25 SQOs = 24.0%, cost/SQO = $4,000, pipeline = $72,000, pipe/$ = $3.00. Either way linkedin_ads is the weakest paid channel; the inversion doesn't change the ranking.

## Recommendation

Shift budget from linkedin_ads into paid_search, with a smaller test allocation to webinars. linkedin_ads is worst on every paid efficiency metric (highest cost/SQO at $3,000 vs $2,000 paid_search and $1,800 webinars; lowest conversion at 32.0%, or 24.0% after excluding the two inverted rows; lowest pipe/$ at $4.00 — $3.00 excluding inversions — vs $20.00 paid_search). One caveat before cutting deeply: paid_search's advantage is partly deal-size-driven — pipe/SQO is $40,000 for paid_search vs $12,000 linkedin_ads and webinars and $9,000 organic — so the reallocation bets that paid_search's larger deals are reproducible, which 18 SQOs supports more strongly than any other channel's sample. paid_social ($18,000, zero tracked first-touch contacts) should be paused or fixed for attribution first; six months of untracked spend is the most clearly waste-like line item, but "zero attributed" is not proof it drives no value — confirm tracking before concluding. Suggested direction: move ~$15–18k of the $24k linkedin_ads budget into paid_search, keep $3–5k webinars-level spend as a diversification test, hold paid_social at $0 pending attribution fix.

## Confidence

Moderate on paid_search, low-to-moderate everywhere else. SQO counts per channel are small: paid_search 18, organic_search 10, linkedin_ads 8, referral 6, webinars 5. With n=8, linkedin_ads' 32% rate has a 95% CI of roughly 15–54%, so it may not truly differ from webinars or organic — the paid_search-vs-linkedin gap is the only comparison with meaningful sample support. Additionally, SQO lags run up to ~22 days after SQM, so August-dated SQMs (8 rows across channels) may not have fully matured, slightly deflating late-window conversion rates. The recommendation direction is robust; the magnitude of the shift is not — re-evaluate after one quarter of reallocated spend.
competitor-battlecard-updatePASS 1.00
Plain-English question: Update the Rivally battlecard with sourced pricing and where each side wins.
Prompt — what the model saw
Update the battlecard for the competitor. Sections: one-line positioning, pricing with source and date (newer source wins, note the conflict), where they win, where we win, objections and responses, recent changes, our 12-month win/loss record against them. Cite a snippet id for every factual claim. Rep opinion on a call is not a fact about the competitor. Anything from the old card you cannot re-source gets marked unverified.
cost $0.0313 · 123s · in 21,307 / out 15,854 tok
✓ competitor_named✓ pricing✓ win_loss✓ unverified_marked✓ citations

✓ Correct Answer (ground truth)

Plain-English answer: Rivally at $7 per user per month (the August source wins over the stale $5, conflict noted), a 13-7 win/loss record, and two unverified old claims flagged.
The answer in one line: competitor = Rivally
citation_ids_available25 items — e.g. S01, S02, S03, S04
competitorRivally
pricingvalue_per_user_per_month=7; source_snippet=S17; source_date=2026-08-12
rep_opinion_snippet_ids_not_factsS09, S21
unverified_old_claimsRivally lacks a Slack integration., Rivally was acquired by WorkHuman in 2025.
win_loss_12mowins=13; losses=7
graded checkscompetitor_named · pricing · win_loss · unverified_marked · citations

Model's actual answer

show the model's full answer
# Battlecard: Rivally (updated 2026-09-25)

**One-line positioning:** Points-based recognition feed for mid-market, expanding into EU enterprise with multi-language support and EU data residency [S02, S04, S12, S15].

**Pricing (newer source wins; conflict noted):**
- Current: Recognition Starter = $7/user/month, annual billing required — pricing page, 2026-08-12 [S17]. Corroborated by deal quote: $7/user/mo list, 15% discount offered for 3-year term → $5.95 effective (7 × 0.85) — call notes, 2026-08-14 [S18].
- Conflict: earlier pricing pages showed $5/user/mo annual (2026-01-20 [S03]; still $5 on 2026-04-01 [S08]). Superseded by date: +$2 = +40% increase between April and August 2026. A June deal quote of $6.50/user/mo (500-seat, annual) [S13] sits between the two list prices and is a deal quote, not list.
- Rivally Pulse survey add-on is priced separately, not bundled — 2026-09-01 [S23]; launched 2026-03-05 [S06].
- "Aggressive discounting" as a general pattern is AE opinion (Elena Sinclair) and NOT confirmed [S21]; confirmed discounting is the single 15%/3-yr offer [S18].

**Where they win:**
- EU: data residency pitched in deals 2026-02-18 [S05], GA + Dublin office 2026-07-01 [S15]; strong for distributed EU teams, multi-language praised [S12]; ex-Workday VP EMEA hired to lead EU expansion [S11].
- Engaging recognition feed [S02, S16]; fast setup (<1 week) with Slack integration out of the box [S04]; support response under 4 hours [S22]; Microsoft Teams app v2 in public preview [S19].

**Where we win:**
- Analytics depth: 800-seat prospect chose Bonusly over Rivally citing analytics [S25]; Rivally analytics described as limited [S02] and dashboards basic vs enterprise tools [S07]; CSV-only exports make exiting Rivally hard [S20].
- Enterprise admin: no SCIM provisioning, manual user management [S10]; no bulk recognition editing [S24]; admin tooling lags peers [S16].
- Head-to-head record is ours: 13W-7L over the last 12 months (65%), 3-0 since June 2026.

**Objections and responses:**
- "Rivally is cheaper" — List is $7 [S17]; their 15% cut requires a 3-year lock [S18] ($5.95 effective), and leaving is costly because exports are CSV-only [S20]. You give up analytics depth [S25, S02, S07].
- "Rivally has EU data residency" — True, GA since 2026-07-01 [S15]; but EMEA rewards catalog is thinner than their US catalog [S14].
- "Rivally sets up faster / support is responsive" — Acknowledge both [S04, S22]; pivot to analytics and admin scale (SCIM, bulk editing) [S10, S24, S25].

**Recent changes:**
- 2026-09-01: Pulse add-on exits beta, priced as add-on [S23]
- 2026-08-20: Teams app v2 public preview [S19]
- 2026-08-14: 15% discount for 3-year term observed in deal [S18]
- 2026-08-12: Starter price $5 → $7 (+40%) [S17]
- 2026-07-01: Dublin office; EU data residency GA [S15]
- 2026-05-09: ex-Workday VP EMEA hired [S11]

**Our 12-month win/loss record vs Rivally (deals file, 2025-09 → 2026-08):**
- Total: 20 deals — 13 wins, 7 losses = 65% (13/20).
- Monthly: 2025-09: 1W-1L; 2025-10: 2W-0L; 2025-11: 1W-1L; 2025-12: 1W-1L; 2026-01: 2W-0L; 2026-02: 2W-0L; 2026-03: 1W-1L; 2026-04: 0W-2L; 2026-05: 0W-1L; 2026-06: 1W-0L; 2026-07: 1W-0L; 2026-08: 1W-0L.
- Trend: 75% (9/12) Sep 2025–Feb 2026 vs 50% (4/8) Mar–Aug 2026; 4 of 7 losses clustered Mar–May 2026 (Deal-5645A5, Deal-C6FFAA, Deal-9066A6, Deal-72A02F); 3-0 since June 2026.
- Note: the deals file has no loss-reason field; the only sourced loss reason is analytics depth (S25, 2026-09-03). AE opinion that Rivally's UI is clunky [S09] is unverified and excluded from competitor facts.

**Old-card items marked UNVERIFIED:**
- "Rivally lacks a Slack integration" — unverified; contradicted by S04 (Slack integration worked out of the box).
- "Rivally was acquired by WorkHuman in 2025" — unverified; no snippet supports it (S11 is a hire from Workday, not an acquisition).
- "Pricing starts at $5" — superseded by S17 (2026-08-12), see Pricing above.
nurture-sequence-diagnosisPASS 1.00
Plain-English question: What is broken in each email sequence — rates, overlaps, tracking errors?
Prompt — what the model saw
Review the sequences. Per sequence: sent, open rate, reply rate, meeting rate, weakest step. Identify tracking errors (opened above sent), audience overlap between sequences, and the failure mode for anything under 2 percent reply. One change per weak sequence and which to fix first. Under 350 words.
cost $0.0602 · 154s · in 33,799 / out 27,907 tok
✓ tracking_error✓ overlap_pair✓ weakest_step✓ fix_priority

✓ Correct Answer (ground truth)

Plain-English answer: Per-sequence open, reply, and meeting rates with the weakest step (all three weakest at step 3), an opened-above-sent tracking error, and audience overlap flagged.
The answer in one line: fix_priority_first = Cold Outbound - HR Leaders
all_pairwise_overlaps['Cold Outbound - HR Leaders', 'Cold Outbound - People Ops'], ['Expansion Nurture', 'New Logo Nurture']
audience_overlap_pairCold Outbound - HR Leaders, Cold Outbound - People Ops
failure_modeCold Outbound - HR Leaders=opens healthy (~37%) but replies near zero from step 1 -> body/CTA not landing, not a deliverability or subject-line problem
fix_priority_firstCold Outbound - HR Leaders
per_sequence{"New Logo Nurture": {"sent": 1386, "opened": 490, "replied": 90, "meetings": 27, "open_rate": 0.3535, "reply_rate": 0.0649, "meeting_rate": 0.0195, "weakest_step": 3}, "Expansion Nurture": {"sent": 8
tracking_errorsequence=Expansion Nurture; step=2; sent=300; opened=340; issue=opened exceeds sent
under_2pct_replyCold Outbound - HR Leaders
graded checkstracking_error · overlap_pair · weakest_step · fix_priority

Model's actual answer

show the model's full answer
SEQUENCE SCORECARD (all rates on total sent)

- Cold Outbound - HR Leaders — sent 1,785; open 545/1,785 = 30.5%; reply 8/1,785 = 0.4%; meetings 0/1,785 = 0.0%. Weakest step: 3 (reply 1/590 = 0.2%).
- Cold Outbound - People Ops — sent 1,163; open 340/1,163 = 29.2%; reply 29/1,163 = 2.5%; meetings 6/1,163 = 0.5%. Weakest step: 3 (6/377 = 1.6%).
- New Logo Nurture — sent 1,386; open 490/1,386 = 35.4%; reply 90/1,386 = 6.5%; meetings 27/1,386 = 1.9%. Weakest step: 3 (18/428 = 4.2%).
- Expansion Nurture — sent 875; open 565/875 = 64.6%*; reply 59/875 = 6.7%; meetings 12/875 = 1.4%. Weakest step: 3 (12/275 = 4.4%).

TRACKING ERROR (opened > sent)
- Expansion Nurture step 2: 340 opened vs 300 sent (+40, 113.3%). Mechanically impossible — logging/suppression defect; all step-2 opens for this sequence are unusable until fixed.

AUDIENCE OVERLAP (audiences.csv: 963 rows, 940 unique contacts)
- 21 contacts sit in both Cold Outbound - HR Leaders and Cold Outbound - People Ops (CT-000849, CT-000884, CT-000890, CT-000908, CT-001033, CT-001097, CT-001101, CT-001103, CT-001105, CT-001130, CT-001153, CT-001159, CT-001217, CT-001227, CT-001236, CT-001255, CT-001258, CT-001277, CT-001285, CT-001311, CT-001345) — 9.7% of People Ops' audience double-touched.
- 2 contacts in both Expansion and New Logo Nurture (CT-000301, CT-000624).

FAILURE MODES (reply < 2%)
- HR Leaders (0.4%): opens look normal but reply-per-open is 8/545 = 1.5% — messages opened, never answered: offer/ICP misfit, worsening each step (0.83% → 0.34% → 0.17%); zero meetings on 1,785 sends. Open figures may also be inflated by auto-openers.
- People Ops step 3 (1.6%): reply-per-open is healthy (29/340 = 8.5%); the failure is low opens (29.2% sequence, 21.2% at step 3) — deliverability/subject-line problem plus step-3 fatigue.

ONE CHANGE EACH
- HR Leaders: cut to step 1 only; steps 2–3 yielded 3 replies on 1,185 sends. Rebuild targeting/offer before re-adding touches.
- People Ops: replace step-3 reminder with a fresh value-add hook to arrest open decay.
- New Logo: rewrite step 3 around a customer proof point (reply falls 8.4% → 6.6% → 4.2% by step).
- Expansion: fix the step-2 open-count logging error, then retest step 3.

FIX FIRST: HR Leaders — 46% of all sends at 0.4% reply and zero meetings.
weekly-marketing-goals-rollupPASS 1.00
Plain-English question: Are we on pace on the quarter's marketing goals?
Prompt — what the model saw
Produce the weekly marketing goals update for the active quarter: SQMs, SQOs, DS2s, closed-lost MIA rate, same-quarter close count, active pipeline coverage against target. For each: QTD actual, target, delta, pace (on/behind/ahead given days elapsed). One paragraph on what moved this week.
cost $0.0124 · 52s · in 5,976 / out 5,343 tok
✓ metrics_present✓ deltas✓ coverage_ratio✓ pace✓ paragraph

✓ Correct Answer (ground truth)

Plain-English answer: SQMs ahead of pace, SQOs on pace, DS2s behind (35 short of expected at day 66 of 92), with deltas and same-quarter closes.
The answer in one line: days_in_quarter = 92; days_elapsed = 66; coverage_ratio_active_pipeline = 0.75; elapsed_fraction = 0.7174 (+1 more below)
days_in_quarter92
coverage_ratio_active_pipeline0.75
days_elapsed66
elapsed_fraction0.7174
metrics{"SQMs": {"qtd_actual": 230, "target": 300, "delta": -70, "expected_at_day_elapsed": 215.22, "pace": "ahead"}, "SQOs": {"qtd_actual": 84, "target": 120, "delta": -36, "expected_at_day_elapsed": 86.09,
pace_ruleexpected = target * days_elapsed/days_in_quarter; within +/-5% of expected = on, above = ahead, below = behind; lower_better metrics invert.
this_week_paragraph_required1
graded checksmetrics_present · deltas · coverage_ratio · pace · paragraph

Model's actual answer

show the model's full answer
# Weekly Marketing Goals Update — Q3-2026 (day 66 of 92)

Pace benchmark: 66/92 = 71.7% of quarter elapsed. All "expected @ pace" figures = target × 0.7174.

**SQMs**
- QTD actual: 230 | Target: 300 | Delta: −70 (76.7% attained)
- Expected @ pace: 300 × 66/92 = 215.2 → actual +14.8 ahead
- Pace: **AHEAD** (current run rate 230/66 = 3.48/day; needs only 70/26 = 2.69/day to hit target)

**SQOs**
- QTD actual: 84 | Target: 120 | Delta: −36 (70.0% attained)
- Expected @ pace: 120 × 66/92 = 86.1 → actual −2.1 vs pace line
- Pace: **MARGINALLY BEHIND** (needs 36/26 = 1.38/day vs current 1.27/day, ~9% acceleration)

**DS2s**
- QTD actual: 40 | Target: 75 | Delta: −35 (53.3% attained)
- Expected @ pace: 75 × 66/92 = 53.8 → actual −13.8 vs pace line
- Pace: **BEHIND** (needs 35/26 = 1.35/day vs current 40/66 = 0.61/day, ~2.2× current run rate)

**Closed-lost MIA rate** (lower is better)
- QTD actual: 5/25 = 20.0% | Target: ≤10% | Delta: +10 pts over ceiling
- Pace: **BEHIND** — rate is not time-phased; at 2× the ceiling it fails outright. To finish ≤10%, need ≥25 further non-MIA losses with zero additional MIA (5/50 = 10%), or proportionally fewer MIA tags.

**Same-quarter closes**
- QTD actual: 10 | Target: 20 | Delta: −10 (50.0% attained)
- Expected @ pace: 20 × 66/92 = 14.3 → actual −4.3 vs pace line
- Pace: **BEHIND** (needs 10/26 = 0.38/day vs current 0.15/day, ~2.5× current run rate)

**Active pipeline coverage**
- QTD actual: $3,000,000 | Target: $4,000,000 | Delta: −$1,000,000 (75.0% coverage)
- Expected @ pace: $4,000,000 × 0.7174 = $2,869,565 → actual +$130,435 ahead
- Pace: **MARGINALLY AHEAD** (needs $38.5K/day vs current $45.5K/day run rate — on trajectory)

**What moved this week**

Not computable from the data provided — these CSVs are a single QTD snapshot with no prior-week baseline, so week-over-week deltas for any metric cannot be derived and I won't invent them. What the QTD standing does show: the conversion-side and late-stage metrics are the drag — DS2s (53.3% attainment vs 71.7% pace) and same-quarter closes (50.0%) require roughly 2.2–2.5× their current run rates over the remaining 26 days, and the MIA rate at 20% is double its ceiling. Volume metrics (SQMs) and pipeline dollars are tracking ahead of pace, so the shortfall pattern points at mid-to-late funnel progression and loss tagging rather than top-of-funnel generation. If you can supply last week's snapshot, I'll produce the true weekly movement paragraph.
communication 5 tests
ceo-slack-compression0.80
Plain-English question: Compress the forecast analysis into one Slack message the CEO can read in ten seconds.
Prompt — what the model saw
Compress the forecast analysis below into a single Slack message for the CEO. Rules: at most three sentences; under 80 words; lead with the decision or the so-what; keep every distinct number that matters and the one caveat that changes the conclusion; cut any sentence that does not carry a claim, a number, or a directive. No em dashes, no throat-clearing phrases.
cost $0.0044 · 22s · in 607 / out 1,980 tok
✓ max_3_sentences✓ under_80_words✓ weighted_present✗ caveat_present✓ no_em_dash

✓ Correct Answer (ground truth)

Plain-English answer: Three sentences, under 80 words, decision first, every distinct number kept plus the one caveat — all checked by deterministic text rules.
The answer in one line: weighted_number = 115976; must_keep = data-quality caveat — deleting the bad news fails
weighted_number115,976
limitmax 3 sentences, max 80 words, no em-dashes
must_keepdata-quality caveat — deleting the bad news fails
graded checksmax_3_sentences · under_80_words · weighted_present · caveat_present · no_em_dash

Model's actual answer

show the model's full answer
Q3 weighted forecast is $115,976.75 (7 COMMIT deals = $44,729; 35% of $203,565 BEST_CASE), but it is understated: 32 deals worth $227,575 carry October close dates, including COMMIT Deal-D348E1 at $13,770. Direct reps to re-qualify those 32 deals before quarter end; dates are being pushed, not re-qualified. Treat the $116k as directional only; the extract looks precise but is not.
stalled-deal-followup-emailPASS 1.00
Plain-English question: Write a follow-up email for this stalled deal — under 120 words, one prior touch, one ask.
Prompt — what the model saw
Write a follow-up email for this stalled deal. Under 120 words. Reference one specific prior touch from the engagement history, make one clear ask, and match the tone of the prior thread. No invented details.
cost $0.0043 · 23s · in 416 / out 1,957 tok
✓ word_cap✓ references_touch✓ single_ask

✓ Correct Answer (ground truth)

Plain-English answer: An email that references the July 28 milestone-awards demo, makes one clear ask, and invents nothing — Deal-0D2F7A, 31 days since last engagement.
The answer in one line: amount = 5100; word_cap = 120; days_since_last_engagement = 31; deal_alias = Deal-0D2F7A (+4 more below)
amount5,100
days_since_last_engagement31
word_cap120
allowed_facts14 items — e.g. Deal-0D2F7A, $5,100, DS3, Alex Franklin
allowed_newENG-00001, ENG-00002, ENG-00003, ENG-00004, ENG-00005
deal_aliasDeal-0D2F7A
must_referenceengagement_id=ENG-00002; date=2026-07-28; type=meeting; why=most substantive prior touch; concrete prospect reactions to milestone awards and points catalog
ownerAlex Franklin
single_cta1
stageDS3
tonewarm, concise, professional (matches prior thread)
graded checksword_cap · references_touch · single_ask

Model's actual answer

show the model's full answer
Subject: Re: 150-seat pricing

Hi,

Following up on my August 5 email with the demo recap and pricing for 150 seats. At the July 28 demo, the People team reacted strongly to automated milestone awards and the points catalog, so I wanted to make sure the pricing reached the right person.

Would a quick call to walk through the 150-seat pricing be helpful? Happy to resend the recap as well.

Best,
Alex Franklin

(71 words. Note: no recipient name appears in the data, so the greeting is generic — the July 28 attendees are listed only as "VP People plus two HR admins.")
funnel-digest-paragraphPASS 1.00
Plain-English question: Write the weekly GTM digest — one paragraph per team, every number traceable.
Prompt — what the model saw
Write the weekly GTM digest: one paragraph per team (Marketing, Sales, CS, PLG) in Bonusly brand voice. Every number must trace to the extracts. No customer contact names, no employee performance commentary, no compensation. Four paragraphs, under 400 words total.
cost $0.0449 · 88s · in 37,514 / out 6,065 tok
✓ four_paragraphs✓ word_cap✓ teams

✓ Correct Answer (ground truth)

Plain-English answer: Four paragraphs (Marketing, Sales, CS, PLG) using only the allowed numbers, no contact names, no performance commentary, under 400 words.
The answer in one line: word_cap = 400
word_cap400
allowed_numbers16 items — e.g. 2, 3, 9, 08
bannedcustomer contact names, employee performance commentary, compensation
required_paragraphsMarketing, Sales, CS, PLG
graded checksfour_paragraphs · word_cap · teams

Model's actual answer

show the model's full answer
# Weekly GTM digest: week of 2026-08-31

**Marketing.** Marketing delivered 46 SQMs against a weekly target of 52, a 6 SQM shortfall that works out to 88.5% attainment (46 ÷ 52 = 0.8846). Webinars were the largest contributor with 18 of the 46 SQMs, 39.1% of the week's total (18 ÷ 46 = 0.3913). Paid search cost per SQM held at $150, so channel efficiency stayed flat while volume came in under plan.

**Sales.** Sales converted 14 SQOs and set 9 DS2 meetings. New pipeline created came to $310,000, roughly $22,143 per SQO on a straight weekly split ($310,000 ÷ 14 = $22,142.86). Three deals closed in-quarter during the week (same-quarter close count: 3). The extract gives no SQO, meeting, or pipeline targets, so attainment against goal can't be computed.

**CS.** Customer success saved 2 renewals and moved team NPS to 61. The extract provides no prior NPS reading, so the size and direction of that move can't be quantified. 3 red-flag accounts remain open heading into next week.

**PLG.** PLG added 412 new signups with activation at 31%, about 128 activated users (412 × 0.31 = 127.72). 38 companies reached the aha moment of 10 recognition gives. Signups count users while the aha-moment count covers companies, different units, so no ratio between the two is computed.
partner-digest-sectionPASS 1.00
Plain-English question: Write the partner digest so every partner in the roster appears exactly once.
Prompt — what the model saw
Write the partner digest for the period. Every partner in the roster appears once: active plays get status with numbers, quiet partners get one line. Partner-sourced pipeline count and amount must match the deal data. Nothing omitted, nothing added.
cost $0.0038 · 14s · in 467 / out 1,489 tok
✓ all_partners✓ pipeline_amount✓ quiet_one_line

✓ Correct Answer (ground truth)

Plain-English answer: Four active partners with status and numbers, four quiet partners with one line each, and partner-sourced pipeline that reconciles to the deal data.
The answer in one line: rule = every roster partner named exactly once; quiet partners get one line; no partners outside roster.
active_no_dealsWorkWell Group
partner_sourced_pipelinecount=5; amount=550000
per_partner_pipeline{"Apex Rewards Co": {"count": 2, "amount": 275000}, "HRCloud Partners": {"count": 1, "amount": 140000}, "CultureBridge": {"count": 2, "amount": 135000}}
quiet_partnersRecogniq, KudosWave, PeopleFirst Advisors, TotalPerk
rosterpartner_name=Apex Rewards Co; status=active | partner_name=HRCloud Partners; status=active | partner_name=CultureBridge; status=active | partner_name=WorkWell Group; status=active | partner_name=Recogniq; status=quiet | partner_name=KudosWave; status=quiet … (+2 more)
ruleevery roster partner named exactly once; quiet partners get one line; no partners outside roster.
graded checksall_partners · pipeline_amount · quiet_one_line

Model's actual answer

show the model's full answer
PARTNER DIGEST — 2026-08-24 to 2026-09-04

Partner-sourced pipeline: 5 deals, $550,000
Arithmetic: 180,000 + 95,000 + 140,000 + 60,000 + 75,000 = 550,000

ACTIVE

Apex Rewards Co
- 2 sourced deals, $275,000 (Deal-DDAAF2 $180,000 + Deal-2CF33E $95,000); Slack reports both in DS1, UTM Source = Partner. Co-webinar locked for 09-15.

HRCloud Partners
- 1 sourced deal, $140,000 (Deal-F1CDA5); integration referral cleared security review and moved to DS2.

CultureBridge
- 2 sourced deals, $135,000 (Deal-096E1D $60,000 + Deal-067213 $75,000); reported early stage from the lunch-and-learn.

WorkWell Group
- 0 sourced deals this period. Requested Q4 restart of the joint playbook; planning call booked 09-09.

QUIET

- Recogniq: no activity or sourced deals reported this period.
- KudosWave: no activity or sourced deals reported this period.
- PeopleFirst Advisors: no activity or sourced deals reported this period.
- TotalPerk: no activity or sourced deals reported this period.

Cross-check: Slack-reported sourced opps (Apex 2 + HRCloud 1 + CultureBridge 2) = 5, matching the 5 deal rows; totals tie to $550,000. Not in data: pipeline by stage beyond Slack notes, and any activity for the four quiet partners.
rfp-answer-from-policyPASS 1.00
Plain-English question: Answer the 10 RFP items from the policy excerpts only — no guessing.
Prompt — what the model saw
Answer the 10 questionnaire items from the policy excerpts only. Cite the excerpt id for every answer. Where the excerpts do not answer the question, mark it unanswerable and say what document would; do not guess. Tone: vendor response, direct.
cost $0.0030 · 15s · in 467 / out 883 tok
✓ answered_count✓ unanswerable_flagged✓ citations✓ no_guess

✓ Correct Answer (ground truth)

Plain-English answer: Seven answered with excerpt citations (Q1-Q7), three flagged unanswerable with the document that would answer them (Q8-Q10).
The answer in one line: answered_count = 7; rule = unanswerable items must be flagged and the document that would answer them named; do not guess.
answered_count7
answer_map{"Q1": ["P1"], "Q2": ["P1"], "Q3": ["P2"], "Q4": ["P3"], "Q5": ["P4"], "Q6": ["P5"], "Q7": ["P6"]}
ruleunanswerable items must be flagged and the document that would answer them named; do not guess.
unanswerableQ8, Q9, Q10
valid_citation_idsP1, P2, P3, P4, P5, P6, P7, P8
graded checksanswered_count · unanswerable_flagged · citations · no_guess

Model's actual answer

show the model's full answer
R1 — Security Questionnaire Responses (source: policy_excerpts.md only)

Q1. Encrypted at rest? Yes — customer data is encrypted at rest using AES-256. [P1]

Q2. Encrypted in transit? Yes — TLS 1.2 or higher. [P1]

Q3. SOC 2? Yes — a SOC 2 Type II report is maintained and renewed annually; a copy is available to prospects under NDA. [P2]

Q4. Retention after termination: Contract term + 90 days. After that, data is deleted from production systems. Arithmetic: retention window = contract duration + 90 days. [P3]

Q5. Subprocessors: Current list published at trust.bonusly.com; customers receive 30 days notice before a new subprocessor is added. [P4]

Q6. SAML SSO: Yes — SAML 2.0 single sign-on, available on the Enterprise plan. (SCIM provisioning is also available on that plan.) [P5]

Q7. Backups: Production databases are backed up daily; backups are encrypted and retained for 35 days. [P6]

Q8. ISO 27001: Unanswerable from the excerpts. The excerpts cover only SOC 2 Type II [P2]; no excerpt addresses ISO 27001. The certifications/compliance overview page on the trust portal (e.g., trust.bonusly.com certifications listing) would confirm whether an ISO 27001 certificate is held.

Q9. Uptime SLA percentage: Unanswerable from the excerpts. No excerpt states an SLA. The contractual uptime SLA would be found in the Master Service Agreement / SLA exhibit or order form.

Q10. HIPAA BAA: Unanswerable from the excerpts. No excerpt addresses HIPAA or Business Associate Agreements. The legal/contracting documents (HIPAA compliance policy or standard BAA template from the legal team) would answer this.
ops-maintenance 5 tests
skill-file-reconciliation0.86
Plain-English question: Audit a skill set against its manifest and flag every conflict with a severity and an action.
Prompt — what the model saw
Reconcile the skill set below against its manifest. Find and report (with severity CRITICAL/WARNING/INFO and an action type TRIM_DESC, UPDATE_BODY, DELETE_SKILL, MERGE, or REVIEW): (1) any two skills whose ALWAYS-trigger phrases overlap/duplicate, (2) any circular delegation chain between skills (name it), (3) any delegation target that does not exist (dangling, name it), (4) any version conflict and which skill should survive, (5) how many manifest descriptions exceed 1,024 characters, (6) any hardcoded page ids, dates, or person names in skill bodies, (7) manifest drift in BOTH directions (files with no manifest row, and manifest rows with no file). Make one proposal per finding; rewrite nothing.
cost $0.1059 · 352s · in 72,191 / out 49,486 tok
✗ duplicate_trigger_found✓ circular_chain✓ dangling_target✓ oversized_desc_count✓ hardcoded_values✓ manifest_drift✓ severity_and_action

✓ Correct Answer (ground truth)

Plain-English answer: A severity-ranked list of TRIM_DESC / UPDATE_BODY / DELETE_SKILL / MERGE / REVIEW actions covering every rule violation found.
The answer in one line: files_audited = 14; rank = CRITICAL / WARNING / INFO with one fix action per issue
files_audited14
findingsduplicate trigger, circular delegation chain, dangling file reference, oversized description, hardcoded values, manifest drift
rankCRITICAL / WARNING / INFO with one fix action per issue
graded checksduplicate_trigger_found · circular_chain · dangling_target · oversized_desc_count · hardcoded_values · manifest_drift · severity_and_action

Model's actual answer

show the model's full answer
# SignalForge Skill-Set Reconciliation

Scope: 14 skill files provided, 14 manifest rows provided. Reconciliation verified programmatically (YAML folded-scalar semantics for descriptions; exact string matching for cross-references). Results below; one proposal per finding; nothing rewritten.

---

## 0. Manifest integrity (context for findings 5 and 7)

- Files = 14, manifest rows = 14. Set difference both directions = empty (arithmetic: 14 − 14 = 0).
- Recomputed every frontmatter description under folded-block-scalar rules (line joins → single spaces, blank line → `\n`, double-quoted scalar unescaping for deal-strategy-coach and email-drafter). All 14 manifest `description_chars` values match measured lengths exactly, e.g. analysis-validator 656 = 656, signalforge-claim-compressor 1006 = 1006. Delta = 0 for all rows.

## 7. Manifest drift — BOTH DIRECTIONS: NONE

- Files with no manifest row: none.
- Manifest rows with no file: none.
- No action required (INFO, no-op). This is the only clean dimension.

---

## 1. ALWAYS-trigger phrase overlap — FOUND

**Finding 1a — CRITICAL — MERGE: comms-drafter vs email-drafter.**
Both descriptions ALWAYS-trigger on identical verbatim phrases: "write me an email," "draft a follow-up," "help me reply" / "help me reply to this," "what should I say," "bump email," "contract nudge," plus the same pasted-message-review trigger. Both bodies cover the same email lanes (outbound, follow-up, post-demo, pricing/stakeholder, contract, EOQ, renewal/expansion, QBR, onboarding) with near-identical tone guidance, and both contain the same contract-follow-up benchmark text (comms-drafter lines ~150; email-drafter §Contract follow-up example). Both claim exclusivity via each other ("this skill drafts, that skill diagnoses" vs "use deal-strategy-coach instead") — but neither excludes the other, so the phrase collision is unresolved.
Proposal: MERGE into a single drafter skill with **email-drafter surviving** (it carries the unique Gmail-signature extraction machinery and is the inbound handoff target from deal-strategy-coach §"Manager-to-prospect email frameworks"); fold comms-drafter's non-email lanes (Intercom/support, partner/channel, tone-anchor table) into it, delete comms-drafter, and update its two inbound references (its own description pointer and the shared "bump email"/"contract nudge" phrases).

**Finding 1b — WARNING — UPDATE_BODY: blanket-ALWAYS trio (model-selection, analysis-validator, signalforge-feedback).**
model-selection: "ALWAYS run this skill at the start of every task, without exception." analysis-validator: "Runs after every SignalForge quantitative analysis… Never skip." signalforge-feedback: "ALWAYS trigger… as the absolute final step… Never skip." On the stated trigger text, every quantitative task fires all three, and the set defines no total ordering except feedback's partial claim ("after analysis-validator and after signalforge-claim-compressor"). model-selection's "before any skill invocation" is compatible but unstated relative to the other two.
Proposal: UPDATE_BODY — add one ordering line to analysis-validator ("runs after model-selection, before signalforge-claim-compressor and signalforge-feedback") so the three blanket triggers form a single declared chain instead of three independent "always" claims. (Overlap between analysis-validator and signalforge-feedback is sequential, not duplicate — no merge needed.)

## 2. Circular delegation chain — FOUND

**Finding 2 — CRITICAL — REVIEW: pipeline-intelligence-report ↔ closed-lost-analysis.**
- pipeline-intelligence-report: Non-Negotiable #5 "Loss Intel tab delegates to closed-lost-analysis skill," and Phase 2b "Delegate entirely to the closed-lost-analysis skill (Mode 4)."
- closed-lost-analysis: Mode 4 "Triggers: … 'flag at-risk deals,' … **called from pipeline-intelligence-report**."
Named chain: pipeline-intelligence-report → closed-lost-analysis (Mode 4) → pipeline-intelligence-report. An executor that follows delegation edges verbatim loops infinitely. Content-wise the split is deliberate (PIR owns orchestration/rendering; CLA Mode 4 returns per-deal loss_risk_score data), so this may be an accepted hub-and-spoke — but it is undocumented as such.
Proposal: REVIEW — add a one-line boundary rule to both bodies ("closed-lost-analysis Mode 4 returns scoring data only and must never re-invoke pipeline-intelligence-report; PIR consumes the return without re-delegating"). Do not merge; the 10-tab report and the standalone loss analysis are legitimately separate outputs.

## 3. Delegation targets not present in the provided set — FOUND

**Finding 3 — WARNING — UPDATE_BODY: 13 skill-name references point outside the reconciled set.**
Existence can only be evaluated against the provided manifest/corpus. Referenced but absent from it (citing skill → target, exactly as given):
- analysis-validator §12.4 → bonusly-data-questions, bonusly-product-questions, bonusly-business-reporting-questions, bonusly-rewards-questions, bonusly-ppp-questions, bonusly-feature-flag-questions, bonusly-deal-desk-questions, bonusly-datadog-questions; §11 → skill-orchestrator
- signalforge-feedback (Activation Checklist) → skill-orchestrator
- comms-drafter → bonusly-brand, prospect-research-multithreading
- email-drafter → bonusly-brand, prospect-research-multithreading
- deal-strategy-coach → prospect-research-multithreading (§Cross-skill handoff)
- signalforge-claim-compressor → bonusly-brand
- sales-forecast → bonusly-brand
- pipeline-intelligence-report → signalforge-reports (referenced only as a `/mnt/skills/organization/…` path, not a manifest skill); same for weekly-pipeline-report
Note: several of these (bonusly-brand, prospect-research-multithreading, the bonusly-*-questions family, skill-orchestrator, signalforge-reports) plausibly exist in the wider Bonusly tree but outside this manifest — if so this is a manifest-coverage gap, not dead links. Also non-skill references with no verification path in the provided data (INFO-grade): `write_to_snowflake` MCP (signalforge-feedback, self-documented as not deployed), `Aligned:get_room_brief` (next-to-close), `ZoomInfo:account_research` (stale-pipeline-report).
Proposal: UPDATE_BODY — either add the 13 missing skills as manifest rows (preferred, since analysis-validator §12.4 hard-depends on 8 of them at runtime) or annotate each reference with its expected location so dangling-vs-external is decidable on the next reconciliation.

## 4. Version conflicts — FOUND

**Finding 4a — CRITICAL — UPDATE_BODY: analysis-validator internal version contradiction. Survivor: v3.6.**
- Header and footer: v3.6; Last Updated May 9, 2026; §7 trail template prints "analysis-validator **v3.2**".
- Changelog rows are out of chronological order (3.6 listed above 3.5) and internally impossible: it jumps "2.6 (May 4, 2026) → 3.6 (May 9, 2026)" and never records the 3.0 restructure the body describes — yet rows for 3.0, 3.1, 3.2, 3.3, 3.4, 3.5 all appear below it, all dated May 9, 2026. There is no 2.7–2.9 and no path from 2.6 to 3.0 in the table's order.
- Which survives: **3.6** — it is the frontmatter/header/footer value, the latest changelog row, and the version other skills cite ("Analysis Validator v3.6" in pipeline-intelligence-report footer).
Proposal: UPDATE_BODY — fix §7 "v3.2" → v3.6 and reorder/rebuild the changelog 1.0 → 3.6 (2.0 → 2.6 → 3.0 → 3.1 → 3.2 → 3.3 → 3.4 → 3.5 → 3.6).

**Finding 4b — WARNING — UPDATE_BODY: pipeline-intelligence-report v6 vs "v4". Survivor: v6.**
Frontmatter `version: v6 · May 2026` and H1 "Full Execution Spec (v6 · May 2026)", but Phase 5 section "## **v4** Component Vocabulary" and footer "✓ SignalForge Validated · Analysis Validator **v3.6**" (that cross-reference is consistent). weekly-pipeline-report separately instructs "using the **v4** design system." Since PIR v6 is the newest declared version and only its section label says v4, v6 survives.
Proposal: UPDATE_BODY — relabel PIR's component-vocabulary section to v6 (or add one line "component vocabulary unchanged from v4"), and update weekly-pipeline-report's "v4 design system" reference to match whatever signalforge-reports currently publishes (not stated in the provided data — flag as unverifiable rather than guessing a number).

## 5. Manifest descriptions exceeding 1,024 characters — ZERO

Arithmetic: 14 descriptions measured; counts 656, 897, 996, 792, 965, 676, 945, 1004, 1006, 962, 1006, 708, 762, 656. Maximum = 1,006 (pipeline-intelligence-report and signalforge-claim-compressor). 1,006 − 1,024 = −18, so none exceed the cap. INFO — no action. Headroom watch: 4 skills sit ≥996 chars (pipeline-intelligence-report 1006, signalforge-claim-compressor 1006, partner-digest 1004, comms-drafter 996); if the MERGE in 1a fires, the combined description must be re-trimmed.

## 6. Hardcoded page IDs, dates, person names in skill bodies — FOUND

**Finding 6a — WARNING — UPDATE_BODY: partner-digest.** Confluence cloud UUID 73fe98de-a4a3-4869-9f8a-bb1eeed4cf7f; page/folder IDs 1958248479, 2286616609 (×4), 2286321666, 2265382925, 2236940297, 2237825028, 2239365136, 2238283777; Slack user ID U03QLMBL7AR; person names Amani Phipps (owner line ×2), Kelli, Jen Lee, Hani, Bryce, Sara; dates May 16, 2026 / May 19, 2026 / June 2, 2026 / 2026-05-17 (changelog). Partially mitigating: the skill itself mandates "update this list as programs change."
Proposal: UPDATE_BODY — move IDs/URLs and the partner-contact roster into a `references/` file the skill reads at run time (keep one canonical copy), same pattern sales-forecast already uses.

**Finding 6b — WARNING — UPDATE_BODY: sales-forecast.** Space ID 2232811524, parent page ID 2232582148, cloud UUID 73fe98de-…; persons Alaina (VP Sales, ×2), Elena (changelog, as the removed hardcode); date July 9, 2026; hardcoded quarter framing: title says "always use the active quarter, not Q2" yet contains "Q2 deals analyzed" footer template, "Q2 Narrative" tab name, and Q2 examples.
Proposal: UPDATE_BODY — parameterize the tab/footer quarter token, move Confluence IDs to references/.

**Finding 6c — WARNING — UPDATE_BODY: signalforge-feedback.** Page IDs 2295136266, 2234417154, 2247295002, spaceId 2232811524, cloud UUID 73fe98de-…, plus full Atlassian URL.
Proposal: UPDATE_BODY — same references/-file extraction.

**Finding 6d — WARNING — UPDATE_BODY: deal-strategy-coach.** Confluence playbook URL bonusly1612893911.atlassian.net/…pages/2257879045 ("April 2026" in title); person-name routing: ".edu … routed to **Farid** for manual qualification," India "routed to **Perseus**"; pricing labeled "2026".
Proposal: UPDATE_BODY — replace named-person routing with role-based routing ("workspace-routing owner, see roster reference"); keep the roster in one shared reference file.

**Finding 6e — WARNING — UPDATE_BODY: weekly-pipeline-report.** Person in the H1 itself ("# Weekly Pipeline Report — **Ben Lavin** · Demand Generation"), Google spreadsheet IDs 1CLZeOsElVDF_LF0ZG_t2nfwvhnZ6bpwqM_nX3WEYzcw and 1ENuaEcCuLjdKhMvp8FK3Ys1ek5Aw9ZuOZhsHJJFoB_k, hardcoded Q2 window "April 1 – June 30, 2026" and static Q1 2026 figures $365,152 / $2,490,532 inside the body (self-contradicting its own "always read live" rule).
Proposal: UPDATE_BODY — replace the person title with role title, move spreadsheet IDs + static historicals to references/queries.md.

**Finding 6f — WARNING — UPDATE_BODY: pipeline-intelligence-report.** Five AE owner IDs with names (Bryce Harmon 119337721, Dana Mercer 83155923, Cole Ingram 83155924, Alex Franklin 84342457, Gavin Porter 1520255671, "verified May 2026"), HubSpot org ID 1973303, stage IDs 150582536/150582537/150582538/150582539/1175632767, /mnt/skills file paths. This directly contradicts stale-pipeline-report's doctrine: "**Never hardcode rep names or owner IDs.** The AE roster changes" (which correctly resolves owners dynamically and hardcodes only channel C0561C1JCPJ — flagged here as its one token; date 2026-06-10 in changelog is acceptable changelog usage).
Proposal: UPDATE_BODY — port stale-pipeline-report's Phase 2 dynamic owner resolution into pipeline-intelligence-report Phase 1; keep stage IDs/pipeline 'default' (those are system constants, defensible); add a REVIEW to standardize one owner-resolution convention across the set (currently three: dynamic [stale-pipeline-report], cached roster [sales-forecast references], hardcoded [PIR]).

**Finding 6g — INFO — REVIEW: analysis-validator.** Full GTM roster with persons + HubSpot owner IDs (§12.3: Amani Phipps 210200121, John Thomas 78303262, Yasmin Wahid 89062643, Alaina Loori 82535637, Shealagh Coughlin 119069206, Ben Castelli 348210196, 6 AEs, 7 CSMs; "Updated May 4, 2026"), person escalation "Manish or Amani" (G1-K ×2, §10), dates March 28, 2023 / May 4, 2026 / April 26, 2026 / May 9, 2026, population anchors ~452,000 / ~110,097, stage IDs. Mostly defensible: it self-labels anchors as session-refresh calibration and stage IDs as canonical tables. The roster is the exposure — it duplicates the same IDs pipeline-intelligence-report hardcodes, so a roster change breaks two files.
Proposal: REVIEW (low urgency) — consider making §12.3 a live `HubSpot:search_owners` pull with the static roster as fallback; stage-ID tables stay as-is.

**Clean on hardcodes:** comms-drafter, email-drafter, closed-lost-analysis (dates only in the defensible "confirmed May 2026" field-verification sense), model-selection body — except model-selection's registry `last_checked: 2026-05-19` is a self-expiring hardcoded date (its own rule: stale after 14 days, so it is guaranteed stale at any later run). **Finding 6h — INFO — UPDATE_BODY:** add to model-selection a run-time check that triggers the self-update procedure on first use rather than relying on the stored date.

---

## Summary counts

| # | Finding | Severity | Action |
|---|---|---|---|
| 1a | comms-drafter ↔ email-drafter ALWAYS-phrase duplication | CRITICAL | MERGE (email-drafter survives) |
| 1b | blanket-ALWAYS trio, no declared ordering | WARNING | UPDATE_BODY |
| 2 | pipeline-intelligence-report ↔ closed-lost-analysis circle | CRITICAL | REVIEW |
| 3 | 13 skill refs outside the set (incl. skill-orchestrator ×2, bonusly-brand ×4, prospect-research-multithreading ×3) | WARNING | UPDATE_BODY |
| 4a | analysis-validator v3.6 vs §7 "v3.2" + impossible changelog | CRITICAL | UPDATE_BODY (v3.6 survives) |
| 4b | pipeline-intelligence-report v6 vs "v4 Component Vocabulary" | WARNING | UPDATE_BODY (v6 survives) |
| 5 | descriptions > 1,024 chars | — | none (max 1,006) |
| 6a–6h | hardcoded IDs/dates/persons (partner-digest, sales-forecast, signalforge-feedback, deal-strategy-coach, weekly-pipeline-report, pipeline-intelligence-report, analysis-validator, model-selection) | WARNING ×6, INFO ×2 | UPDATE_BODY ×7, REVIEW ×1 |
| 7 | manifest drift both directions | — | none (14 = 14) |

Totals: 14 findings — CRITICAL 3, WARNING 9, INFO 2. Actions: MERGE 1, REVIEW 3, UPDATE_BODY 8, no-op 2. Nothing was rewritten.
runbook-from-slack-threadPASS 1.00
Plain-English question: Turn the resolved incident thread into a runbook.
Prompt — what the model saw
Turn the resolved incident thread into a runbook: ordered steps with the exact command or action taken, who ran it, how success was verified, and rollback for each step that changed state. Every step must trace to a message in the thread. Add nothing the thread does not contain; where the thread is unclear, mark the step as needing confirmation.
cost $0.0066 · 26s · in 1,609 / out 2,038 tok
✓ ordered_steps✓ rollback✓ needs_confirmation✓ sourced

✓ Correct Answer (ground truth)

Plain-English answer: Ordered steps from the thread, each with the exact command, who ran it, how success was verified, and rollback where state changed.
The answer in one line: rule = no steps beyond what the thread contains; every step traces to a message id.
ambiguous_steporder=4; msg=M05; mark=needs-confirmation
ruleno steps beyond what the thread contains; every step traces to a message id.
stepsorder=1; msg=M01; actor=Bryce Harmon; action=acknowledge PagerDuty alert, take IC; state_change=False; verified_by=M01 | order=2; msg=M02; actor=Farid Osman; action=check queue depth; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M02 (48,213 pending) | order=3; msg=M04; actor=Farid Osman; action=pause auto-enqueue via feature flag; command=bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'; state_change=True; rollback=bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'; verified_by=M07 (queue depth falling) | order=4; msg=M05; actor=Elena Sinclair; action=cleared the dead set; state_change=True; needs_confirmation=thread gives no command or method for clearing the dead set | order=5; msg=M06; actor=Bryce Harmon; action=scale reward-worker to 6 replicas; command=kubectl scale deployment/reward-worker --replicas=6; state_change=True; rollback=kubectl scale deployment/reward-worker --replicas=3; verified_by=M07/M08 | order=6; msg=M08; actor=Cole Ingram; action=verify queue drained and error rate baseline; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M08 (returns 0) … (+2 more)
graded checksordered_steps · rollback · needs_confirmation · sourced

Model's actual answer

show the model's full answer
Runbook below, built only from the supplied thread. Each step cites its source message ID.

REWARD-WORKER QUEUE INCIDENT RUNBOOK
Source: #incident-2026-08-29-reward-queue thread export (resolved). Actors, commands, and numbers taken exactly as recorded.

STEP 1 — Acknowledge alert and take IC [M01]
- Action: Bryce Harmon acknowledged the PagerDuty alert (reward-worker queue depth > 10k) and took IC.
- Who: Bryce Harmon (2026-08-29 14:02:10Z)
- Command: none recorded (PagerDuty acknowledge).
- Success verification: none documented — needs confirmation.
- Rollback: N/A (acknowledgment/role assumption; no state changed).

STEP 2 — Measure queue depth [M02]
- Action: Measured reward queue depth.
- Who: Farid Osman (14:04:33Z)
- Command: `bundle exec rake sidekiq:queue_depth`
- Result reported: 48,213 pending jobs; normal is under 500.
- Success verification: command output itself (read-only diagnostic).
- Rollback: N/A (read-only).

STEP 3 — Inspect dead set [M03]
- Action: Inspected Sidekiq dead set contents.
- Who: Farid Osman (14:06:02Z)
- Command: not recorded in thread — needs confirmation.
- Result reported: 112 dead jobs, all Redis::TimeoutError, timestamped around 13:58.
- Success verification: command output itself (read-only inspection); inspection command absent — needs confirmation.
- Rollback: N/A (read-only inspection; no state changed in this step).

STEP 4 — Stop the bleed: disable auto-recognition enqueue [M04]
- Action: Disabled the enqueue feature flag.
- Who: Farid Osman (14:08:45Z)
- Command: `bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'`
- Success verification: no direct flag-state verification recorded — needs confirmation.
- Rollback (documented in M04): `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`

STEP 5 — Clear the dead set [M05]
- Action: Cleared the Sidekiq dead set.
- Who: Elena Sinclair (14:15:20Z), via the console.
- Command: not recorded — needs confirmation.
- Success verification: none documented — needs confirmation.
- Rollback: not documented — needs confirmation. (This is a destructive, state-changing action. Nothing in the thread establishes it as an approved repeatable remediation; treat as one-time action pending confirmation.)

STEP 6 — Scale reward-worker deployment 3 → 6 [M06]
- Action: Scaled reward-worker replicas up from 3 to 6.
- Who: Bryce Harmon (14:21:07Z)
- Command: `kubectl scale deployment/reward-worker --replicas=6` (was 3)
- Success verification: no direct replica-count check recorded. M07 and M08 are later queue observations, not isolated proof of this step — needs confirmation.
- Rollback (documented in M06): `kubectl scale deployment/reward-worker --replicas=3`

STEP 7 — Interim observation (no state change) [M07]
- Action: Observed drain rate.
- Who: Farid Osman (14:33:41Z)
- Result reported: queue depth down to 9,400, falling ~1,200/min.
- Command for the measurement: not recorded — needs confirmation.
- Rollback: N/A (observation only).

STEP 8 — Verify recovery [M08]
- Action: Verified queue drained and error rate normalized.
- Who: Cole Ingram (14:47:55Z)
- Command: `bundle exec rake sidekiq:queue_depth`
- Result reported: returns 0; Datadog error rate back to baseline.
- Success verification: command output plus Datadog error-rate check (as reported).

STEP 9 — Re-enable enqueue [M09]
- Action: Re-enabled the enqueue feature flag.
- Who: Bryce Harmon (14:49:10Z)
- Command: `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`
- Success verification: 40 new jobs processed cleanly in the next 3 minutes (verified by the actor).
- Rollback: not documented — needs confirmation. (Do not infer Step 4's disable command as the approved rollback for this step; it is not documented as such in the thread.)

STEP 10 — Scale reward-worker deployment back to 3 [M10]
- Action: Scaled reward-worker replicas back down to 3.
- Who: Bryce Harmon (14:55:00Z)
- Command: `kubectl scale deployment/reward-worker --replicas=3`
- Success verification: queue stable at 0; incident declared resolved.
- Rollback: not documented — needs confirmation. (Do not infer scaling back to 6 as rollback; it is not documented as such in the thread.)

NEEDS-CONFIRMATION SUMMARY
- Step 3: dead-set inspection command absent.
- Step 4: no direct flag-state verification recorded.
- Step 5: command, verification, and rollback all absent (destructive action — must not be repeated without confirmation).
- Step 6: no direct replica-count verification recorded.
- Step 7: measurement command for the 9,400 / ~1,200/min figures absent.
- Steps 9–10: no documented rollback.
datadog-incident-timelinePASS 1.00
Plain-English question: Sequence the failure from the logs — what broke first, and what cascaded after it?
Prompt — what the model saw
From the log slice, sequence the failure: first error with timestamp, the cascade in order, the service and job involved, and the Datadog query that would confirm the first error. State what the logs do not show.
cost $0.0068 · 34s · in 1,668 / out 3,550 tok
✓ first_error_ts✓ service_named✓ cascade_ordered✓ query✓ gaps_stated

✓ Correct Answer (ground truth)

Plain-English answer: A reward-service Redis timeout at 14:01:12Z cascades to sidekiq at 14:01:40 and api-gateway at 14:03:05, with the Datadog query that confirms the first error.
The answer in one line: job_class_involved = RewardGiveJob
cascade_orderservice=reward-service; first_error_at=2026-09-03T14:01:12Z | service=sidekiq; first_error_at=2026-09-03T14:01:40Z | service=api-gateway; first_error_at=2026-09-03T14:03:05Z | service=web-app; first_error_at=2026-09-03T14:03:30Z
clean_servicespostgres
confirming_queryservice:reward-service status:error "@message:Redis::TimeoutError"
first_errortimestamp=2026-09-03T14:01:12Z; service=reward-service; message=Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s
job_class_involvedRewardGiveJob
logs_do_not_showroot cause of the Redis timeout, redis-server's own logs, infrastructure metrics (CPU/memory/network), any deploy or config change preceding the incident, who applied mitigation or when (no actor fields in logs)
graded checksfirst_error_ts · service_named · cascade_ordered · query · gaps_stated

Model's actual answer

show the model's full answer
**1) First error**

2026-09-03T14:01:12Z — reward-service, ERROR:
"Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s"

Last healthy signal before it: 13:59:30Z (reward-service INFO "job enqueued") — a silent gap of 1m42s precedes the first error. The error log is the first observed symptom, not necessarily the failure onset.

**2) Cascade in order**

1. 14:01:12Z reward-service — Redis connection timeout (first error)
2. 14:01:20 / 14:01:30 / 14:01:40Z reward-service — 3× "retry exhausted for RewardGiveJob"
3. 14:01:40Z sidekiq — RewardGiveJob failed; "retrying in 60s"; then repeated failures at 14:02:28, 14:03:31, 14:04:22, 14:05:26, 14:06:47Z (6 failure lines total)
4. 14:02:30Z sidekiq WARN — queue reward depth above 10,000
5. 14:02:36Z onward (jobs table) — blast radius widens to a second job class: RecognitionDigestJob J-00013–J-00016 fail (14:02:36–14:05:50Z); RewardGiveJob J-00001–J-00012 fail (14:01:46–14:02:58Z). Total 16 failed jobs = 12 + 4
6. 14:03:05Z api-gateway — first 502: "upstream timeout calling reward-service /gives" (1m53s after first error); further 502s at 14:03:48, 14:04:13, 14:05:16, 14:06:52Z (5 total)
7. 14:03:30Z web-app — first customer-visible failure: "Give form submission failed: upstream 502 from api-gateway" (2m18s after first error); further at 14:04:45, 14:05:42, 14:06:49Z (4 total)
8. Errors cease after 14:06:52Z; silent window of 15m18s (only postgres checkpoints)
9. 14:22:10Z reward-service — "Redis connection restored; resuming job processing" (20m58s after first error)
10. 14:24:45Z sidekiq — queue reward depth below 500 (23m33s after first error; 2m35s after restore)

**3) Service and job involved**

- Failing dependency: Redis at redis-primary:6379 (5s connection timeout)
- Origin service: reward-service (Sidekiq-backed)
- Primary job: RewardGiveJob (J-00001–J-00012); collateral: RecognitionDigestJob (J-00013–J-00016)
- Propagation path: Redis → reward-service → sidekiq (queue depth >10,000) → api-gateway (502s) → web-app (give form failures)

**4) Datadog query to confirm the first error**

service:reward-service status:error "Redis::TimeoutError"

Time range 2026-09-03 13:55:00–14:05:00 UTC, sorted oldest-first; the first event returned should be the 14:01:12Z line with "timed out after 5s". Narrower variant: service:reward-service status:error "redis-primary:6379 timed out after 5s".

**5) What the logs do not show**

- No redis-primary logs at all — no cause of the Redis failure (failover, memory pressure, network, restart) and no record of when Redis actually went down.
- No cause of recovery — nothing between 14:06:52Z and 14:22:10Z except postgres checkpoints; "Redis connection restored" has no preceding action logged.
- Root-cause linkage from job_id to log line — the earliest job-table failure (J-00005, 14:01:46Z) post-dates the first error (14:01:12Z), so no failed job_id can be tied to the first timeout.
- Total job volume — jobs table lists failures only; succeeded jobs are uncounted, so no failure rate can be computed.
- Customer/tenant impact — no customer, user, or deal identifiers on any log line or job record; affected-entity count is unknown.
- Queue depth between the two bounds — only ">10,000" at 14:02:30Z and "<500" at 14:24:45Z; no drain rate or backlog peak value.
- The sidekiq log line at 14:02:28Z ("retrying", 48s after the 14:01:40Z "retrying in 60s" line) does not reconcile with a 60s retry interval; the logs don't explain the discrepancy.
- Any alerting, on-call, deploy, or change events around 14:01Z.
feature-flag-state-summaryPASS 1.00
Plain-English question: Summarize what each feature flag controls and who has it on.
Prompt — what the model saw
Summarize the state of the listed feature flags: what each controls per the code excerpt, which companies or segments have it on, targeting rules, and any flag with no code reference or no targets. Company counts must match the export. Do not describe a flag that is not in the export.
cost $0.0075 · 42s · in 494 / out 4,437 tok
✓ orphans✓ no_targets✓ counts_match✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: Recognition streaks on for 42 beta companies, redeem-flow redesign off with a 12-company target list, plus any flag with no code reference or no targets.
The answer in one line: rule = company counts must equal the export; do not describe flags absent from the export.
code_referencedrecognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, redeem_flow_redesign, analytics_dashboard_v3, ms_teams_app_v2
flags{"recognition_streaks_v2": {"state": "on", "targeting_rules": "segment:beta_companies", "company_count": 42}, "points_budget_guardrails": {"state": "on", "targeting_rules": "all_companies", "company_c
flags_with_no_targetspaused_offboard_cleanup
orphan_flags_no_code_referencelegacy_give_modal, survey_boosters_q3, paused_offboard_cleanup
rulecompany counts must equal the export; do not describe flags absent from the export.
graded checksorphans · no_targets · counts_match · no_invented

Model's actual answer

show the model's full answer
FLAG STATE SUMMARY — 9 flags in export; 6 with code references, 3 without

FLAGS WITH CODE REFERENCES

1. recognition_streaks_v2 — ON, segment:beta_companies, 42 companies
   app/models/recognition.rb: when enabled, StreakTracker.record(give) runs per give (streak tracking). No else branch shown, so off-path behavior isn't in the excerpt.

2. points_budget_guardrails — ON, all_companies, 220 companies
   app/services/budget_service.rb: when enabled, BudgetService.enforce!(giver, points) — enforces points budgets on gives. Count 220 with all_companies targeting implies a 220-company universe.

3. slack_dm_nudges — ON, segment:region_na, 87 companies
   app/jobs/nudge_job.rb: gates SlackDm.send_nudge(user); job returns early when disabled (Slack DM nudges).

4. redeem_flow_redesign — OFF, targeted_list, 12 companies in export
   app/controllers/redeem_controller.rb: toggles RedeemV2Component (on) vs RedeemV1Component (off).
   Caveat: state is OFF yet the export counts 12 companies under the targeted list. The export doesn't define whether company_count means actively served or merely configured targets, so whether those 12 actually get V2 cannot be determined from the data.

5. analytics_dashboard_v3 — ON, segment:tier_three, 65 companies
   app/controllers/analytics_controller.rb: when enabled, @dashboard = AnalyticsV3.new(company). No else branch shown — off-path dashboard behavior not in excerpt.

6. ms_teams_app_v2 — OFF, targeted_list, 9 companies in export
   app/services/teams_installer.rb: when enabled, TeamsAppV2.install(company) (legacy install path not shown).
   Same caveat as flag 4: OFF state + 9 targeted companies; active-vs-configured ambiguity unresolved by the export.

FLAGS WITH NO CODE REFERENCE IN flag_code.md — what they control cannot be determined from the provided excerpt

7. legacy_give_modal — OFF, segment:legacy_plan, 14 companies. No code reference found; off per export, so nothing is served regardless.
8. survey_boosters_q3 — ON, segment:legacy_plan, 7 companies. No code reference found: enabled for 7 legacy_plan companies, but the effect is unknown from the excerpt. Note: shares the legacy_plan segment with legacy_give_modal yet shows 7 vs 14 companies; the export does not explain the delta (legacy_plan universe is at least 14 on this data).
9. paused_offboard_cleanup — OFF, no targeting rules, 0 companies. No code reference, no targets. Fully dormant per the export.

ROLLOUT MATH (denominator = 220, implied by all_companies on flag 2)
- recognition_streaks_v2: 42/220 = 19.1%
- points_budget_guardrails: 220/220 = 100% (fully rolled out)
- slack_dm_nudges: 87/220 = 39.5%
- analytics_dashboard_v3: 65/220 = 29.5%
- survey_boosters_q3: 7/220 = 3.2%

TALLY: 5 ON / 4 OFF. Segments referenced: beta_companies, region_na, tier_three, legacy_plan; plus all_companies (1 flag) and targeted_list (2 flags: redeem_flow_redesign, ms_teams_app_v2). Every flag in the code excerpt maps to a flag in the export — no orphan code references.

MISSING/AMBIGUOUS, EXPLICITLY: (a) company_count semantics undefined for the two OFF targeted_list flags; (b) no code for flags 7–9; (c) no off-path branches for flags 1 and 5; (d) legacy_plan 14-vs-7 delta unexplained.
nda-triagePASS 1.00
Plain-English question: Triage each NDA green, yellow, or red — and quote the clause that decides it.
Prompt — what the model saw
Triage each NDA as GREEN (standard approval), YELLOW (counsel review), or RED (full legal review). Cite the clause that drives each classification by quoting it. Catch embedded non-solicits, non-competes, missing carve-outs, and unusual term or governing law. One line of reasoning per NDA.
cost $0.0030 · 12s · in 547 / out 859 tok
✓ green✓ yellow✓ red✓ non_solicit_caught✓ clause_quoted✓ correct_labels

✓ Correct Answer (ground truth)

Plain-English answer: NDA-1 green (standard mutual terms), NDA-2 yellow for an embedded 18-month non-solicit, with the driving clause quoted for every classification.
The answer in one line: non_solicit_in = NDA-2
checksclassification_per_nda, clause_quotes_exist_in_document, non_solicit_caught, carveouts_checked
classifications{"NDA-1": {"classification": "GREEN", "driving_reason": "standard mutual NDA: 2-year term, standard carve-outs, Delaware law", "driving_clause_quote": "Carve-outs: information that (a) is or becomes p
non_solicit_inNDA-2
graded checksgreen · yellow · red · non_solicit_caught · clause_quoted · correct_labels

Model's actual answer

show the model's full answer
NDA-1 — [PARTY A] — GREEN (standard approval)
Driving clause: "(a) is or becomes publicly available through no breach, (b) was known prior to disclosure, (c) is received from a third party without duty of confidence, (d) is independently developed, or (e) must be disclosed by law or court order" — all five standard carve-outs present, plus "No license, no obligation to proceed, no exclusivity"; mutual, 2-yr term/3-yr survival and Delaware law are routine.

NDA-2 — [PARTY B] — YELLOW (counsel review)
Driving clause: "During the term of this Agreement and for eighteen (18) months thereafter, neither party shall solicit for employment or hire any employee of the other party with whom it came into contact under this Agreement" — embedded non-solicit (18 months post-term) layered onto otherwise standard terms (mutual, standard carve-outs, Delaware); it is mutual and contact-scoped, so not auto-RED, but counsel should review scope before signature.

NDA-3 — [PARTY C] — RED (full legal review)
Driving clauses: "For a period of three (3) years following the Effective Date, Recipient shall not, directly or indirectly, engage in or provide services to any business that competes with Discloser's business" and "No confidentiality carve-outs are stated in this Agreement" — an unusually broad 3-year non-compete with no competing-business definition, zero carve-outs, one-way confidentiality on a long 5-year term, plus unusual foreign governing law ("laws of the Republic of Ireland, and the parties submit to the exclusive jurisdiction of its courts"); compounding terms require full legal review.