← back to all models ← RevenueBench home
openai

openai/gpt-6-astra

Mean 0.952 · 34/40 perfect tests · $19.82 total · modalities in: text, image, file · out: text · each card: the prompt → the correct answer (gold) → this model's actual answer

What do the modalities mean?

Modalities in means what you can send this model: text, images, files, video, or audio. Modalities out means what it can send back. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what this model can handle in real work.

data-crm 5 tests
stage-hygiene-auditPASS 1.00
Plain-English question: How much of the open DS2+ pipeline is missing a substantive why-buys, and is anything missing DS1 timestamps or out of order?
Prompt — what the model saw
Audit stage hygiene on the open pipeline below. Rules: a deal at DS2 or later must have a substantive why-buys (15+ characters, not a placeholder such as tbd, n/a, see notes). A why_buys value of the form LONG(n) means the real text was present but truncated for display; treat it as present and substantive. Every deal should have a DS1 entry timestamp. A deal whose stage entry timestamps are out of order regressed (an earlier-stage timestamp is later than a later-stage timestamp). Report: percent of DS2+ deals non-compliant on why-buys overall and by owner (use the owners table for names), count of deals with no DS1 timestamp, count that regressed, and three process fixes. State definitions before numbers.
cost $0.7630 · 84s · in 9 / out 5,359 tok
✓ noncomp_pct✓ no_ds1_0✓ regressed_0✓ per_owner_breakdown✓ definitions_first✓ three_fixes

✓ Correct Answer (ground truth)

Plain-English answer: 81.5% of the 124 DS2+ deals are missing a substantive why-buys — Alex Franklin 48, Cole Ingram 19, Dana Mercer 18, Bryce Harmon 15, Farid Osman 1 — and every deal has a DS1 timestamp with none regressed.
The answer in one line: n_deals = 156; ds2plus = 124; noncomp_pct = 81.5; bryce_noncomp = 15; no_ds1 = 0; regressed = 0 (+1 more below)
n_deals156
bryce_noncomp15
ds2plus124
no_ds10
noncomp_pct81.5
regressed0
per_owner_noncompAlex Franklin=48 of 61 (78.7%); Cole Ingram=19 of 22 (86.4%); Dana Mercer=18 of 22 (81.8%); Bryce Harmon=15 of 18 (83.3%); Farid Osman=1 of 1 (100%)
graded checksnoncomp_pct · no_ds1_0 · regressed_0 · per_owner_breakdown · definitions_first · three_fixes

Model's actual answer

show the model's full answer
Definitions
- Why-buys non-compliance: DS2–DS5 deals with blank, fewer than 15 characters, or placeholder why_buys. LONG(n) is compliant. Denominator: DS2+ deals, overall or per owner.
- Missing DS1: blank t_ds1 across all open deals.
- Regressed: any populated earlier-stage timestamp later than a populated later-stage timestamp. Equal dates are allowed; missing timestamps cannot establish regression.

Why-buys non-compliance
Overall: 156 open − 32 DS1 = 124 DS2+; 124 − 23 compliant = 101 non-compliant.
101 / 124 × 100 = 81.45%.

Owner                  Non-compliant / DS2+ × 100
Bryce Harmon             15 / 18 × 100 = 83.33%
Dana Mercer             18 / 22 × 100 = 81.82%
Alex Franklin          48 / 61 × 100 = 78.69%
Cole Ingram      19 / 22 × 100 = 86.36%
Farid Osman             1 /  1 × 100 = 100.00%
Elena Sinclair           0 /  0 = N/A; no DS2+ deals

Gavin Porter and Hugo Lindqvist have no open deals; rates are N/A.

Timestamp hygiene
- Missing DS1: 156 − 156 populated = 0 deals.
- Regressed: 0 deals with timestamp inversions. Missing later-stage timestamps limit detection.

Three process fixes
1. Gate DS2+ entry and updates on substantive why-buys: enforce minimum length, reject placeholders, and require owners to remediate existing blanks.
2. Automatically capture DS1 entry timestamps; backfill missing values only from verified stage history.
3. Preserve append-only stage-transition history and flag timestamp inversions for review without overwriting genuine regressions.
owner-resolutionPASS 1.00
Plain-English question: Every open deal has a blank owner name — resolve each one to a real person using the owners table.
Prompt — what the model saw
The deal extract below has an owner name field that is blank on nearly every row and an owner id field that is populated. Using the owners table, resolve each open deal to an owner name. Report: number of deals resolved, deals whose owner id has no match in the owners table (list them), the owner ids that map to archived/deactivated owners (list their names), and the total pipeline amount per resolved owner. Do not guess a name for any unmatched id.
cost $0.6162 · 53s · in 9 / out 3,006 tok
✓ resolved_count✓ unmatched_none✓ deactivated_flagged✓ jake_amount✓ justin_amount✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: All 156 deals resolve to six named owners; jake_amount maps to Bryce Harmon ($1,054,144) and justin_amount to Alex Franklin ($624,310).
The answer in one line: jake_amount = 1,054,144; justin_amount = 624,310; resolved = 156
jake_amount1,054,144
justin_amount624,310
resolved156
deactivated
per_ownerBryce Harmon=1,054,144; Alex Franklin=624,310; Dana Mercer=341,195; Cole Ingram=288,161.43; Farid Osman=4,134; Elena Sinclair=2,100
graded checksresolved_count · unmatched_none · deactivated_flagged · jake_amount · justin_amount · no_invented

Model's actual answer

show the model's full answer
Resolved: 156 of 156 open deals, using exact owner_id matches.
Arithmetic: 35 + 24 + 67 + 22 + 7 + 1 = 156.

Unmatched deals: None.

Archived owners in owners.csv:
- 1520255671 — Gavin Porter
- 77260721 — Hugo Lindqvist

Neither archived owner has any deals in the supplied extract. All open deals map to owners with archived=false; no separate deactivation field was provided.

Total pipeline per resolved owner (all stages and forecast categories; currency not specified):

| Owner ID | Resolved owner | Deals | Pipeline amount |
|---|---|---:|---:|
| 119337721 | Bryce Harmon | 35 | 1,054,144.00 |
| 83155923 | Dana Mercer | 24 | 341,195.00 |
| 84342457 | Alex Franklin | 67 | 624,310.00 |
| 83155924 | Cole Ingram | 22 | 288,161.43 |
| 716654662 | Farid Osman | 7 | 4,134.00 |
| 701163055 | Elena Sinclair | 1 | 2,100.00 |
| Total | | 156 | 2,314,044.43 |

Arithmetic by owner, using stage subtotals in DS1 + DS2 + DS3 + DS4 + DS5 order:
- Bryce Harmon: 645,497 + 168,091 + 176,400 + 50,656 + 13,500 = 1,054,144.
- Dana Mercer: 32,000 + 146,975 + 105,750 + 37,470 + 19,000 = 341,195.
- Alex Franklin: 53,516 + 197,800 + 303,690 + 23,574 + 45,730 = 624,310.
- Cole Ingram: 0 + 187,152.35 + 96,919.08 + 2,760 + 1,330 = 288,161.43.
- Farid Osman: 2,885 + 1,249 + 0 + 0 + 0 = 4,134.
- Elena Sinclair: 2,100 + 0 + 0 + 0 + 0 = 2,100.

Total: 1,054,144 + 341,195 + 624,310 + 288,161.43 + 4,134 + 2,100 = 2,314,044.43.
gong-hubspot-join-integrityPASS 1.00
Plain-English question: How many call records point at deal ids that do not exist in the open pipeline?
Prompt — what the model saw
Check the integrity of the call-to-deal mapping below. The gong table lists calls in the last 90 days keyed to HubSpot deal ids; deals_open lists the current open default-pipeline deals. Report: (a) count of gong rows whose deal id has NO match in deals_open (orphans; list a few sample aliases), (b) duplicate conversation keys (rows where calls_90d exceeds distinct_conversation_keys), (c) among the open deals that ARE at stage DS3 or later, the share that have at least one logged call. Give sample aliases for the orphan finding, all taken from the data. Note: gong_owner_id is a Gong user id, not a HubSpot owner id.
cost $0.7876 · 34s · in 15 / out 997 tok
✓ orphan_count✓ dup_keys_0✓ ds3_coverage✓ sample_ids_present

✓ Correct Answer (ground truth)

Plain-English answer: 30 Gong call rows are orphans keyed to dead deal ids, none are duplicates, and DS3 call coverage is 29.4%.
The answer in one line: orphans = 30; ds3_cov_pct = 29.4; dup = 0
orphans30
ds3_cov_pct29.4
dup0
orphan_aliasesDeal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42, Deal-36EA09
graded checksorphan_count · dup_keys_0 · ds3_coverage · sample_ids_present

Model's actual answer

show the model's full answer
(a) Orphans: 67 Gong rows − 37 matching deals_open.deal_id = 30 orphan rows.
Sample aliases: Deal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42. These are unmatched in the supplied open-deal table, not necessarily invalid deals.

(b) Duplicate conversation keys: 0 rows where calls_90d > distinct_conversation_keys. All 67 rows have equal counts; total excess = 465 − 465 = 0. Individual conversation keys were not provided, so cross-deal duplication cannot be checked.

(c) DS3+ call coverage, using current stage:
- Eligible open deals: 61 DS3 + 14 DS4 + 10 DS5 = 85.
- With at least one logged call: 12 DS3 + 6 DS4 + 7 DS5 = 25.
- Share: 25 ÷ 85 × 100 = 29.41%.

Joined on hs_deal_id = deal_id; gong_owner_id was not used.
snowflake-sql-trapped-catalog0.50
Plain-English question: Write one Snowflake query for first-month usage per customer using only the provided, trap-laden catalog.
Prompt — what the model saw
Using ONLY the catalog provided, write one Snowflake SQL query returning, per customer company, for its first calendar month as a customer: unique givers, recognition count, successful redemption count. Read the catalog notes carefully: some tables are stale or unpopulated and must be avoided, and one documented filter (excluding deleted givers) must NOT be applied to historical giving counts. Put a comment above each FROM or JOIN explaining the table choice. Return only fenced SQL.
cost $0.2187 · 21s · in 3 / out 401 tok
✓ avoids_stale_deals✓ avoids_unpopulated_live✗ uses_cohort_summary✗ succeeded_filter✓ no_giver_deletion_filter✗ returns_sql

✓ Correct Answer (ground truth)

Plain-English answer: A single correct SQL query that uses the right catalog tables and avoids the trap columns — graded on the exact tables and columns.
The answer in one line: use_table = PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
avoid_stalePRODUCTION.HUBSPOT.DEALS (stale)
avoid_unpopulatedV2_LIVE.OBJECTS_DEALS (unpopulated)
forbidden_filterdeleted-giver exclusion must NOT be applied
outputone fenced SELECT with per-FROM/JOIN table-choice comment
required_filtersucceeded redemptions only
use_tablePRODUCTION.PLG.COMPANY_COHORT_SUMMARY
graded checksavoids_stale_deals · avoids_unpopulated_live · uses_cohort_summary · succeeded_filter · no_giver_deletion_filter · returns_sql

Model's actual answer

show the model's full answer
```sql
-- Cannot write the requested query using only this catalog:
-- 1. No customer-company source with a company key and customer-start date
--    is documented, so the first calendar month cannot be determined.
-- 2. No recognition source or company, giver, event ID, and event timestamp
--    columns are documented.
-- 3. PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 is the documented
--    redemption source, but its company key and event timestamp are not given.
--    Its event grain or unique redemption key is also unspecified.
--
-- Required rules once the missing schema is provided:
-- Count historical giving without excluding deleted givers.
-- Count redemptions only where STATE = 'succeeded'.
-- Aggregate giving and redemptions separately to avoid join multiplication.
-- PRODUCTION.PLG.COMPANY_COHORT_SUMMARY covers only self-serve companies;
-- its M1 fields are not documented as first-customer-calendar-month metrics.
-- Do not use unpopulated, stale, or retired tables.
```
crm-completeness-enrichmentPASS 1.00
Plain-English question: Audit the CRM for missing and invalid fields, and propose the fixes.
Prompt — what the model saw
Audit the CRM extract for completeness and propose fixes. Deals need owner, stage, amount, close date, why-buys; companies need industry, employee count, HQ country; contacts need email, title, persona. Report completeness percent per field. Identify duplicate company clusters (name variants, shared domain) with a survivor per cluster. List invalid emails and domain mismatches. Fill missing company fields only where the enrichment export has a matching row; where CRM and enrichment disagree list both and recommend a source. Never invent a value. End with the 10 fixes with the most pipeline amount at stake.
cost $0.6044 · 86s · in 9 / out 5,212 tok
✓ dup_clusters✓ invalid_emails✓ survivor_language✓ completeness_pct✓ no_invent

✓ Correct Answer (ground truth)

Plain-English answer: 9 companies missing employee count, 6 missing HQ country, 4 invalid emails, 1 domain mismatch, 2 duplicate clusters, and 25 enrichment matches.
The answer in one line: n_contacts = 52; n_companies = 34; enrichment_matches = 25; invalid_email_count = 4; n_dup_clusters = 2; domain_mismatch_count = 1 (+4 more below)
n_contacts52
domain_mismatch_count1
enrichment_matches25
invalid_email_count4
n_companies34
n_dup_clusters2
company_missingindustry=0; employee_count=9; hq_country=6
domain_mismatchesCT-0011
duplicate_clusters['C-0A092931', 'C-0A092932'], ['C-0A092933', 'C-0A092934']
invalid_emailsCT-0010, CT-0080, CT-0081, CT-0192
graded checksdup_clusters · invalid_emails · survivor_language · completeness_pct · no_invent

Model's actual answer

show the model's full answer
CRM completeness

Completeness = nonblank values ÷ records × 100, after trimming whitespace. These are pre-fix, pre-deduplication results; populated does not necessarily mean valid.

| Object | Required field | Arithmetic | Completeness |
|---|---|---|---:|
| Deals | Owner | No deal extract supplied | N/A |
| Deals | Stage | No deal extract supplied | N/A |
| Deals | Amount | No deal extract supplied | N/A |
| Deals | Close date | No deal extract supplied | N/A |
| Deals | Why-buys | No deal extract supplied | N/A |
| Companies | Industry | 34 ÷ 34 × 100 | 100.00% |
| Companies | Employee count | 25 ÷ 34 × 100 | 73.53% |
| Companies | HQ country | 28 ÷ 34 × 100 | 82.35% |
| Contacts | Email | 52 ÷ 52 × 100 | 100.00% |
| Contacts | Title | 39 ÷ 52 × 100 | 75.00% |
| Contacts | Persona | 37 ÷ 52 × 100 | 71.15% |

Email syntax validity is separate: (52 − 4) ÷ 52 × 100 = 92.31%. Deliverability cannot be established from these files.

Duplicate company clusters

Company names are not provided, so name-variant detection is unavailable. Shared domains identify these candidate duplicate clusters:

| Shared domain | Members | Proposed survivor | Reconciliation |
|---|---|---|---|
| acme-corp.com | C-0A092931, C-0A092932 | C-0A092931 | Employee count conflicts: 500 versus 510. Industry is `Technology` versus `tech`; country is `US` versus `USA`. |
| globex.io | C-0A092933, C-0A092934 | C-0A092933 | Industry differs: `SaaS` versus `Technology`. Both have employee count 200 and country `US`. |

Survivors use lowest alias as a deterministic tie-breaker—not evidence of accuracy or age. Neither domain has enrichment. Confirm common entity identity before merging; preserve conflicting values and relationship history. Do not treat the survivor’s employee count or industry as verified merely because it survives.

Invalid emails and domain mismatches

| Contact | Company | Email | Issue |
|---|---|---|---|
| CT-0010 | C-66D1FC | `user0@` | Missing email domain |
| CT-0080 | C-92D97D | `user0@` | Missing email domain |
| CT-0081 | C-92D97D | `user1@` | Missing email domain |
| CT-0192 | C-425E2A | `user2@` | Missing email domain |
| CT-0011 | C-66D1FC | `user1@other-domain.com` | Email domain `other-domain.com` differs from company/contact domain `66d1fc.com` |

Verify the invalid addresses; do not construct replacements by appending company domains. Investigate CT-0011’s email and company association rather than automatically changing either. All contact `domain` values match their linked company’s domain.

Supported company fills

Exact domain matches in zoominfo_enrichment.csv support these fills only:

| Company | Matching domain | Field | CRM → enrichment |
|---|---|---|---|
| C-EC3025 | ec3025.com | employee_count | blank → 400 |
| C-96039F | 96039f.com | employee_count | blank → 400 |
| C-44EA29 | 44ea29.com | employee_count | blank → 400 |
| C-D04904 | d04904.com | employee_count | blank → 400 |
| C-B23205 | b23205.com | employee_count | blank → 400 |
| C-60C75F | 60c75f.com | employee_count | blank → 400 |
| C-7BBDFA | 7bbdfa.com | employee_count | blank → 400 |
| C-50D386 | 50d386.com | employee_count | blank → 400 |

After these proposed fills, employee-count completeness becomes (25 + 8) ÷ 34 × 100 = 97.06%, before deduplication.

Leave these fields blank:
- HQ country: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5. Matching enrichment rows also have blank countries.
- HQ country: C-EE9FFB. No matching enrichment row.
- Employee count: C-93C8BF. No matching enrichment row.

No missing HQ country can be filled from the supplied enrichment. A domain alone is not evidence of country.

CRM versus enrichment disagreements

Industry differences:

| Company aliases | CRM industry | Enrichment industry |
|---|---|---|
| C-66D1FC, C-44EA29, C-60C75F | `tech` | `Computer Software` |
| C-EC3025, C-92D97D, C-D04904, C-77A95A, C-AA8DDA, C-B25F40 | `Technology` | `Computer Software` |
| C-425E2A | `Tech `, including trailing space | `Computer Software` |

Recommendation: use enrichment as the proposed source for a detailed industry classification on these matched rows, subject to review. The exports contain no timestamps or verification metadata proving it is more accurate. Preserve the original CRM value and distinguish taxonomy mapping from a factual correction; `Technology` is broader than `Computer Software`.

Country representation differences:

| Company aliases | CRM HQ country | Enrichment HQ country |
|---|---|---|
| C-66D1FC, C-950043, C-77A95A, C-B23205, C-D0662E | `US` | `United States` |
| C-EC3025, C-96039F, C-E51FB7, C-425E2A, C-2D7423 | `USA` | `United States` |

Recommendation: use enrichment’s `United States` as the canonical spelling for these matched records. These are equivalent country representations, not substantive geographic conflicts.

There are no disagreements between populated CRM employee counts and matching enrichment employee counts.

Separately, standardize industry casing, whitespace, and equivalent labels (`tech`/`Tech `/`Technology`, `health care`/`Healthcare`) through an explicit taxonomy. Do not infer `Computer Software` for companies lacking a matching enrichment row.

Missing contact fields

Request verified values for the following; the supplied enrichment contains no contact titles or personas.

- Title: CT-0000, CT-0022, CT-0072, CT-0080, CT-0081, CT-0092, CT-0120, CT-0121, CT-0122, CT-0132, CT-0141, CT-0162, CT-0170.
- Persona: CT-0000, CT-0022, CT-0041, CT-0060, CT-0070, CT-0081, CT-0082, CT-0092, CT-0110, CT-0132, CT-0162, CT-0171, CT-0172, CT-0180, CT-0181.

Do not infer persona solely from title.

10 fixes — pipeline ranking unavailable

No deals, amounts, stages, or deal-to-company/contact associations were supplied. Therefore, the 10 fixes with the most pipeline at stake cannot be identified. The following are unranked fix candidates; pipeline amount is unknown for every item, not zero.

1. Fill the eight supported employee counts with 400 for C-EC3025, C-96039F, C-44EA29, C-D04904, C-B23205, C-60C75F, C-7BBDFA, C-50D386.
2. Verify and correct CT-0010’s invalid email for C-66D1FC.
3. Verify and correct CT-0080 and CT-0081’s invalid emails for C-92D97D.
4. Verify and correct CT-0192’s invalid email for C-425E2A.
5. Resolve CT-0011’s email/company-domain mismatch for C-66D1FC.
6. Validate and merge C-0A092932 into proposed survivor C-0A092931, retaining the unresolved 500/510 employee-count conflict.
7. Validate and merge C-0A092934 into proposed survivor C-0A092933, retaining the unresolved `SaaS`/`Technology` distinction.
8. Obtain verified values for the missing contact titles and personas listed above.
9. Obtain missing HQ countries for C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB, and employee count for C-93C8BF; leave blank until supported.
10. Review the listed CRM/enrichment industry differences and standardize industry and country representations while preserving source values.
deal-intelligence 5 tests
closed-lost-classification0.80
Plain-English question: Why did we lose the last six months of deals — one primary category per deal, and whose side is the reason on?
Prompt — what the model saw
Classify every deal lost in the last 6 months below into one primary category: pricing, competitor, no decision, timing, product gap, champion left, or other. Also assign a side (Bonusly, buyer, unknown). Then summarize: category counts, the side split, the number of deals where the structured closed_lost_tag clearly disagrees with the free-text reason, and the two patterns most worth acting on. Base everything only on the tag and free-text provided.
cost $0.8220 · 160s · in 21 / out 7,092 tok
✓ total_90✗ competitor_top✓ category_counts_present✓ side_split✓ two_patterns

✓ Correct Answer (ground truth)

Plain-English answer: Every closed-lost deal gets exactly one category (pricing, competitor, no decision, timing, product gap, champion left, other) plus a side, per the computed classification.
The answer in one line: total_lost_6mo = 90; top_reason = competitor (26)
total_lost_6mo90
categoriespricing, competitor, no decision, timing, product gap, champion left, other
patternstwo most actionable, tag-vs-text disagreements counted
sideBonusly / buyer / unknown per deal
top_reasoncompetitor (26)
graded checkstotal_90 · competitor_top · category_counts_present · side_split · two_patterns

Model's actual answer

show the model's full answer
Classification rules: use the stated underlying reason before the outcome; retain the tag when text is inconclusive. “Pricing” includes budget constraints. Side identifies the stated source of the obstacle—not blame: Bonusly for offering/value gaps, buyer for internal constraints or preferences, unknown when attribution is unsupported. MIA maps to no decision, but its cause remains unknown.

All supplied deals are included as the requested cohort.

### Deal classifications

| Deal alias | Primary category | Side |
|---|---|---|
| Deal-DB0AAC | timing | buyer |
| Deal-F7F635 | competitor | unknown |
| Deal-AC944F | no decision | unknown |
| Deal-214060 | no decision | unknown |
| Deal-91A056 | timing | buyer |
| Deal-29326C | timing | buyer |
| Deal-5DB9B0 | other | unknown |
| Deal-831B7B | timing | buyer |
| Deal-F97C37 | product gap | Bonusly |
| Deal-13E9CF | no decision | buyer |
| Deal-39E25C | timing | buyer |
| Deal-7ED004 | pricing | buyer |
| Deal-21B045 | no decision | unknown |
| Deal-B3ABED | timing | buyer |
| Deal-422BA6 | competitor | Bonusly |
| Deal-ED9AE7 | other | buyer |
| Deal-988493 | no decision | unknown |
| Deal-381C8C | competitor | unknown |
| Deal-F308CA | no decision | unknown |
| Deal-F1E8A6 | competitor | unknown |
| Deal-B6AC09 | timing | buyer |
| Deal-70F704 | no decision | unknown |
| Deal-E6E80A | timing | buyer |
| Deal-B038F0 | timing | buyer |
| Deal-4664E1 | no decision | unknown |
| Deal-175756 | timing | buyer |
| Deal-E74A73 | no decision | buyer |
| Deal-DDAB52 | product gap | Bonusly |
| Deal-ACE061 | competitor | unknown |
| Deal-BB78F3 | timing | buyer |
| Deal-D48E0B | no decision | unknown |
| Deal-15DA99 | timing | buyer |
| Deal-F4AF5D | timing | buyer |
| Deal-79B7A1 | timing | buyer |
| Deal-583ADB | no decision | unknown |
| Deal-8E27DA | other | buyer |
| Deal-2D2F8D | competitor | unknown |
| Deal-E0441F | no decision | unknown |
| Deal-7CB44D | no decision | unknown |
| Deal-0F96AA | competitor | unknown |
| Deal-1BCA50 | pricing | buyer |
| Deal-7CC678 | competitor | unknown |
| Deal-FAC17C | no decision | buyer |
| Deal-242273 | product gap | Bonusly |
| Deal-50E5D8 | no decision | buyer |
| Deal-A2C349 | competitor | unknown |
| Deal-9F176A | timing | buyer |
| Deal-7B2236 | pricing | Bonusly |
| Deal-AFA56C | no decision | unknown |
| Deal-C7156E | competitor | unknown |
| Deal-C33D91 | pricing | buyer |
| Deal-9048EB | product gap | Bonusly |
| Deal-5E64CE | timing | buyer |
| Deal-8A0992 | competitor | unknown |
| Deal-D0C698 | competitor | buyer |
| Deal-69CF3D | timing | buyer |
| Deal-ECBF89 | timing | buyer |
| Deal-3618CC | product gap | Bonusly |
| Deal-EECC02 | competitor | unknown |
| Deal-5AD03E | competitor | unknown |
| Deal-D1A623 | timing | buyer |
| Deal-413C56 | no decision | buyer |
| Deal-47F1A1 | competitor | buyer |
| Deal-BF2A98 | competitor | buyer |
| Deal-2A292B | no decision | buyer |
| Deal-D1AABF | no decision | unknown |
| Deal-FEDBCB | timing | buyer |
| Deal-1E7DA9 | competitor | unknown |
| Deal-2BBA21 | no decision | unknown |
| Deal-286F9C | competitor | unknown |
| Deal-7FBAC6 | no decision | buyer |
| Deal-369281 | competitor | buyer |
| Deal-386F6E | no decision | unknown |
| Deal-9FCD0D | competitor | buyer |
| Deal-55867E | timing | unknown |
| Deal-DAFB82 | pricing | buyer |
| Deal-2FEDDB | timing | buyer |
| Deal-64B19A | competitor | unknown |
| Deal-3F86A0 | no decision | unknown |
| Deal-096750 | no decision | unknown |
| Deal-F325A5 | no decision | buyer |
| Deal-ABD14C | no decision | unknown |
| Deal-79E61A | no decision | unknown |
| Deal-8A119B | pricing | buyer |
| Deal-AE7C4E | no decision | unknown |
| Deal-DAB4F1 | no decision | unknown |
| Deal-B4B50F | no decision | unknown |
| Deal-981AD4 | product gap | Bonusly |
| Deal-DC77FE | product gap | Bonusly |
| Deal-5885B9 | no decision | unknown |

Classification limitations:
- Deal-ED9AE7 lists “Timing, budget, authroity.” without identifying a primary obstacle; classified other.
- Deal-3618CC’s “Wanted Surveys” is interpreted as a product-fit requirement; the specific deficiency is missing. Deal-5AD03E’s “Wanted more defined budget access” is too ambiguous to establish a product gap, so the competitor tag is retained.
- Vague competitor-tagged reasons do not establish a confirmed competitor selection. HeyTaco in Deal-ACE061 and Motivosity in Deal-64B19A remain explicitly uncertain.
- No reason explicitly establishes that a buyer champion left. Deal-F325A5 mentions leadership change and deprioritization; Deal-E0441F mentions a departed sales rep.

### Category counts and side split

| Category | Bonusly | buyer | unknown | Total |
|---|---:|---:|---:|---:|
| pricing | 1 | 5 | 0 | 6 |
| competitor | 1 | 5 | 15 | 21 |
| no decision | 0 | 8 | 23 | 31 |
| timing | 0 | 21 | 1 | 22 |
| product gap | 7 | 0 | 0 | 7 |
| champion left | 0 | 0 | 0 | 0 |
| other | 0 | 2 | 1 | 3 |
| Total | 9 | 41 | 40 | 90 |

Arithmetic:
- Categories: 6 + 21 + 31 + 22 + 7 + 0 + 3 = 90 deals.
- Sides: 9 Bonusly + 41 buyer + 40 unknown = 90 deals.

### Clear tag–reason disagreements

4 deals = 3 primary-reason mismatches + 1 explicit timing-horizon mismatch. “Lost DM” is interpreted literally as loss of the decision-maker.

| Deal alias | Tag | Free-text disagreement |
|---|---|---|
| Deal-8E27DA | Feature Request | Chose a swag provider and “didn't want R&R”—a different scope, not a stated missing feature. |
| Deal-FAC17C | Lost DM | Could not get approval from the Executive IT Director—not loss of that decision-maker. |
| Deal-3618CC | Lost DM | “Wanted Surveys” describes a product requirement, not decision-maker loss. |
| Deal-9F176A | Lost- Timing (1 year or more) | “Closer to the end of the year” conflicts with the tag’s one-year-or-more horizon. |

Missing corroboration is not counted as clear disagreement. Neither are compatible overlapping reasons: Deal-9048EB explicitly contains both nonresponse and feature gaps; competitor outcomes can coexist with product gaps. “Next year” alone does not establish a one-year-or-more delay.

### Two patterns most worth acting on

1. Timing and non-decisions dominate: 22 + 31 = 53 deals; 53 ÷ 90 × 100 = 58.9%. Separate dated reactivation opportunities from unexplained silence. Use the stated buyer milestones for Deal-91A056, Deal-B3ABED and Deal-5E64CE; use stronger qualification and explicit next-step commitments for nonresponse cases such as Deal-F308CA and Deal-7CB44D. Do not assume those silent deals lost on price or product.

2. Specific offering differences are more actionable than generic competitor labels. Seven product-gap classifications identify broader offerings, customization, surveys, UI/localization or feature fit: Deal-F97C37, Deal-DDAB52, Deal-242273, Deal-9048EB, Deal-3618CC, Deal-981AD4 and Deal-DC77FE. Deal-422BA6 adds an explicit ADP partnership advantage: 7 + 1 = 8 differentiation cases to validate for qualification, positioning, product or partnership action. The notes do not establish the exact missing capabilities in every case.
pipeline-tieringPASS 1.00
Plain-English question: Score all 156 open deals into exactly one of LOCK / ACTION / BUILD / REVIVE / WATCH / RISKY.
Prompt — what the model saw
Score the open pipeline below on these signals: forecast category, stage, meetings_30d, contact count, and engagement recency. Assign exactly one tier per deal: LOCK, ACTION, BUILD, REVIVE, WATCH, or RISKY (RISKY = the forecast category disagrees with the engagement evidence). Never tier a deal LOCK if it has zero meetings_30d. Do NOT list every deal. Return ONLY a compact JSON object with exactly these keys: - "tier_counts": object mapping each tier to the number of deals in it (must sum to the total deal count), - "tier_examples": object mapping each tier to up to 3 example deal aliases from the data, - "risky_deals": array of the deal aliases you tiered RISKY, - "lock_violations": number of deals tiered LOCK that have zero meetings_30d (should be 0), - "pipeline_shape": one short paragraph on the overall pipeline shape. Note: inbound_emails_30d is 0 for every row due to a data defect, so treat meetings_30d as the inbound signal.
cost $0.7953 · 80s · in 12 / out 4,092 tok
✓ json_returned✓ tier_vocabulary✓ tier_counts_sum✓ risky_deals_listed✓ no_lock_violation_claimed

✓ Correct Answer (ground truth)

Plain-English answer: All 156 open deals tiered exactly once across the six tiers using forecast category, stage, meetings, contacts, and recency.
The answer in one line: total = 156
total156
graded checksjson_returned · tier_vocabulary · tier_counts_sum · risky_deals_listed · no_lock_violation_claimed

Model's actual answer

show the model's full answer
{"tier_counts":{"LOCK":3,"ACTION":18,"BUILD":34,"REVIVE":1,"WATCH":65,"RISKY":35},"tier_examples":{"LOCK":["Deal-D348E1","Deal-C26D20","Deal-403845"],"ACTION":["Deal-25F752","Deal-3974EB","Deal-499BF6"],"BUILD":["Deal-944310","Deal-A5E80A","Deal-1FC049"],"REVIVE":["Deal-2D1F1B"],"WATCH":["Deal-6787C2","Deal-66D1FC","Deal-950043"],"RISKY":["Deal-E53952","Deal-5408B0","Deal-9AAE5F"]},"risky_deals":["Deal-E53952","Deal-5408B0","Deal-9AAE5F","Deal-547B2B","Deal-B7EBD1","Deal-A2B47C","Deal-2465CE","Deal-C61CF7","Deal-62D607","Deal-584EE5","Deal-C6D97A","Deal-7B3B0F","Deal-F9A08A","Deal-0660B4","Deal-FD9F4E","Deal-BA571A","Deal-FC22A3","Deal-7BBDFA","Deal-60C2C2","Deal-4A13AD","Deal-8AD4A5","Deal-15D24F","Deal-9D0060","Deal-690476","Deal-635B8E","Deal-ED725A","Deal-55164C","Deal-3BA5EA","Deal-5FDCE4","Deal-F336B6","Deal-5EED42","Deal-BA3DDC","Deal-7599B8","Deal-F9A3C1","Deal-FA32A0"],"lock_violations":0,"pipeline_shape":"3+18+34+1+65+35=156 deals; 55 have meetings, 99 have none, and 2 lack engagement rows (Deal-3EED2C, Deal-57FF13), assigned WATCH. Pipeline is meeting-light: 35 BEST_CASE/COMMIT deals lack the designated inbound signal and are RISKY. Provisional rules use 2026-09-04, the latest supplied email/call date, as the snapshot because none is specified; later meeting dates do not establish recency. LOCK requires COMMIT, DS4–5, at least 3 contacts, positive meetings and engagement within 14 days. ACTION requires the same engagement/contact evidence plus DS3–5 or COMMIT, excluding LOCK; other meeting-positive deals are BUILD. Remaining zero-meeting deals are REVIVE if engagement exceeds 30 days, otherwise WATCH; RISKY takes precedence."}
call-transcript-extractionPASS 1.00
Plain-English question: Extract the CRM write-back fields from each call transcript as JSON.
Prompt — what the model saw
For each transcript, extract CRM write-back fields as JSON: why-buys (prospect statements only), pain points, stakeholders from the speaker list, budget signal (prospect-stated or null), timeline signal, competitor mentioned (only if the prospect raised it), next step (only if explicitly agreed), objections, confidence. Never fill a field from what the rep said or hoped.
cost $0.3162 · 35s · in 6 / out 1,660 tok
✓ json_emitted✓ budget_signal✓ null_discipline✓ stakeholders✓ competitor_prospect_only

✓ Correct Answer (ground truth)

Plain-English answer: Per transcript: prospect-only why-buys, pain points, stakeholders, budget and timeline signals, and competitors — e.g. TX-001: '$40k earmarked', live before open enrollment.
The answer in one line: rep = Alex Franklin
allowed_newTX-001, TX-002, TX-003, TX-004, TX-005, TX-006, Achievers, Bucketlist, Nectar
per_transcript{"TX-001": {"why_buys": ["automating anniversary and birthday awards"], "pain_points": ["HR team of three cannot keep up with it manually", "track everything in a spreadsheet, and people slip through
repAlex Franklin
rep_sourced_nulls{"TX-003": ["budget_signal \u2014 only the rep mentioned pricing/budget"], "TX-004": ["next_step \u2014 rep proposed, prospect said 'no promises'"], "TX-006": ["budget_signal \u2014 rep decoy, prospec
graded checksjson_emitted · budget_signal · null_discipline · stakeholders · competitor_prospect_only

Model's actual answer

show the model's full answer
[
  {
    "transcript_id": "TX-001",
    "deal_alias": "Deal-CFE7F4",
    "why_buys": [
      "Automate anniversary and birthday awards."
    ],
    "pain_points": [
      "HR team of three cannot keep up with awards manually.",
      "Spreadsheet tracking lets people slip through the cracks."
    ],
    "stakeholders": [
      "Prospect (VP People)",
      "Prospect (HR Admin)"
    ],
    "budget_signal": "About $40k earmarked for engagement tools this fiscal year; not stated as exclusively allocated to this purchase.",
    "timeline_signal": "Ideally live before open enrollment in November; year not stated.",
    "competitor_mentioned": [
      {
        "name": "Achievers",
        "context": "Evaluated last year; considered too heavy for a team their size."
      }
    ],
    "next_step": "Security review on September 12, explicitly accepted by Prospect (VP People); year not stated.",
    "objections": [
      "SSO and audit logs required for IT sign-off."
    ],
    "confidence": {
      "extraction": "high",
      "basis": "Direct prospect statements; next step explicitly accepted."
    }
  },
  {
    "transcript_id": "TX-002",
    "deal_alias": "Deal-70BB30",
    "why_buys": [
      "Tie recognition to retention for the hourly workforce."
    ],
    "pain_points": [
      "Regretted turnover in the hourly workforce is over 30%."
    ],
    "stakeholders": [
      "Prospect (Head of Total Rewards)",
      "Prospect (CFO)"
    ],
    "budget_signal": "Finance approved a $25k pilot budget for this quarter.",
    "timeline_signal": "Decision wanted by end of September; year not stated.",
    "competitor_mentioned": null,
    "next_step": "Send the pilot agreement; Prospect (CFO) agreed to route it to legal this week.",
    "objections": [
      "Rock-solid Workday integration is a condition."
    ],
    "confidence": {
      "extraction": "high",
      "basis": "Direct prospect statements; no competitor named; next step explicitly accepted."
    }
  },
  {
    "transcript_id": "TX-003",
    "deal_alias": "Deal-530B50",
    "why_buys": [
      "Make recognition visible across 12 retail locations."
    ],
    "pain_points": [
      "Store managers have zero budget autonomy for on-the-spot recognition."
    ],
    "stakeholders": [
      "Prospect (People Ops Manager)"
    ],
    "budget_signal": null,
    "timeline_signal": "No rush until Q1; year not stated.",
    "competitor_mentioned": [
      {
        "name": "Bucketlist",
        "context": "CEO used it at her last company and liked it; current evaluation not stated."
      }
    ],
    "next_step": "Schedule a call with the CEO; Prospect (People Ops Manager) agreed to send two times.",
    "objections": [
      "CEO must be sold first and decides anything people-related."
    ],
    "confidence": {
      "extraction": "high",
      "basis": "Direct prospect statements; no prospect-stated purchase budget; next step explicitly accepted."
    }
  },
  {
    "transcript_id": "TX-004",
    "deal_alias": "Deal-180D02",
    "why_buys": [
      "Consolidate three separate recognition tools into one."
    ],
    "pain_points": [
      "Paying for three tools, none of which connect to their HRIS."
    ],
    "stakeholders": [
      "Prospect (VP People)",
      "Prospect (IT Security Lead)"
    ],
    "budget_signal": "Prospect (VP People) can approve under $15k annually without board approval; this is an approval threshold, not a committed budget.",
    "timeline_signal": "Procurement takes six to eight weeks minimum; no target purchase or launch date stated.",
    "competitor_mentioned": null,
    "next_step": null,
    "objections": [
      "Security review took three months for the last vendor, creating hesitation."
    ],
    "confidence": {
      "extraction": "high",
      "basis": "Direct prospect statements; no competitors named; proposed CFO follow-up was tentative, not agreed."
    }
  },
  {
    "transcript_id": "TX-005",
    "deal_alias": "Deal-F8767A",
    "why_buys": [
      "Automate service milestones.",
      "Provide analytics on recognition equity across departments."
    ],
    "pain_points": [
      "Night-shift teams feel invisible; their engagement scores run 20 points lower, with the comparison group not specified."
    ],
    "stakeholders": [
      "Prospect (HR Director)",
      "Prospect (People Ops Coordinator)"
    ],
    "budget_signal": "$12k approved under the engagement line.",
    "timeline_signal": "Running before the January all-hands; year not stated.",
    "competitor_mentioned": [
      {
        "name": "Nectar",
        "context": "Currently mid-pilot; the prospect says this offering must beat that experience."
      }
    ],
    "next_step": "Present to the executive team on October 2, explicitly accepted by Prospect (HR Director); year not stated.",
    "objections": [
      "Must beat the current Nectar pilot experience.",
      "Executive team is skeptical after a failed rollout two years ago."
    ],
    "confidence": {
      "extraction": "high",
      "basis": "Direct prospect statements; next step explicitly accepted."
    }
  },
  {
    "transcript_id": "TX-006",
    "deal_alias": "Deal-EE195F",
    "why_buys": [
      "Cut administration time on service awards."
    ],
    "pain_points": [
      "Prospect (HR Manager) personally spends five hours a month ordering and shipping plaques."
    ],
    "stakeholders": [
      "Prospect (HR Manager)"
    ],
    "budget_signal": "“Budget isn't the issue — time is.” No amount or approval stated.",
    "timeline_signal": "Fiscal year resets in February; prospect says a Q1 start is realistic. Year and Q1 calendar basis not stated.",
    "competitor_mentioned": [
      {
        "name": null,
        "context": "No external vendor: prospect is comparing against doing it internally."
      }
    ],
    "next_step": "Send the one-page overview; Prospect (HR Manager) agreed to forward it to the COO this week.",
    "objections": [
      "COO usually prefers building things in-house."
    ],
    "confidence": {
      "extraction": "high",
      "basis": "Direct prospect statements; internal alternative explicitly raised; next step explicitly accepted."
    }
  }
]
next-to-close0.50
Plain-English question: Which 3 deals are closest to signature right now?
Prompt — what the model saw
Which 3 deals are closest to signature right now? Use the CRM data and the Slack thread; Slack often has the real status before the CRM does. For each: deal, amount, why it is close, what is left. Under 250 words.
cost $0.2816 · 23s · in 6 / out 479 tok
✗ exactly_three✓ top3_named✓ slack_first✗ blocked_excluded

✓ Correct Answer (ground truth)

Plain-English answer: Deal-547B2B ($11,200), Deal-B7EBD1 ($9,000), and Deal-403845 ($9,000) — all DS5 COMMIT deals closing September 10-11, with what is left on each.
blocked_excludeddeal_alias=Deal-2465CE; reason=champion left + procurement freeze (Slack #sales-sparring 2026-09-04)
exactly_three1
ranking_ruleSlack signature-imminent deal ranks first; remaining ranked by stage (DS5>DS4), forecast (COMMIT>BEST_CASE>PIPELINE), close_date asc, amount desc
top3rank=1; deal_alias=Deal-547B2B; amount=11200; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11 | rank=2; deal_alias=Deal-B7EBD1; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-10 | rank=3; deal_alias=Deal-403845; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11
graded checksexactly_three · top3_named · slack_first · blocked_excluded

Model's actual answer

show the model's full answer
Best-supported ranking from the CRM and September 4 Slack thread; no later signature confirmations are provided.

1. Deal-547B2B — $11,200
   Why close: CRM shows DS5/COMMIT. Slack says redlines returned clean, the signing page is out, and VP People said they would sign “tomorrow.” Strongest explicit signature signal.
   What is left: Customer signature and confirmation it was received; the stated signing expectation is not proof of completion.

2. Deal-403845 — $9,000
   Why close: CRM shows DS5/COMMIT; Slack confirms the order form is with finance and the deal is “moving fine.”
   What is left: Finance processing/approval and signature; neither is confirmed. Ownership needs reconciliation: CRM lists Alex Franklin, but Dana Mercer reports the update.

3. Deal-A2B47C — $6,360
   Why close: CRM shows DS5/COMMIT with a September 11 close date; Slack confirms it remains warm and is progressing through normal legal review. Provisional third: Deal-D348E1 has the same Slack status but a later CRM close date.
   What is left: Complete legal review and obtain signature; no signing commitment is provided.

Deal-2465CE is excluded despite CRM COMMIT: Slack reports champion departure, a procurement freeze, and a realistic Q4 timeline. Deal-B7EBD1 has no Slack confirmation of readiness.

Amounts are directly reported in next_close_deals.csv; no arithmetic is required.
product-gap-detectionPASS 1.00
Plain-English question: Which prospect complaints are true product gaps versus rollout or enablement issues?
Prompt — what the model saw
Surface true product gaps raised by prospects in active deals. For each candidate: quote the prospect line, classify as real gap, plan gate, or rollout/enablement issue using the product docs excerpt, and name the deal and amount. Only real gaps go in the summary. Rep language does not count as prospect voice.
cost $0.2980 · 19s · in 9 / out 555 tok
✓ real_gaps✓ plan_gate_classified✓ enablement_classified✓ gap_summary

✓ Correct Answer (ground truth)

Plain-English answer: Two real gaps (SCIM provisioning, ADP Workforce Now integration) separated from plan gates and enablement issues, each with the prospect quote and deal.
allowed_newTG-001, TG-002, TG-003, TG-004, TG-005, SCIM user provisioning, ADP Workforce Now integration, Custom report builder
enablement_issuestranscript_id=TG-004; deal_alias=Deal-84DBA6; amount=16000; feature=Slack integration recognition adoption; quote=We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.; why_enablement=Slack integration exists on all plans; managers were never trained
excluded_rep_voicetranscript_id=TG-005; line=Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap.; reason=rep voice, not prospect-raised
plan_gatestranscript_id=TG-003; deal_alias=Deal-CFE7F4; amount=32000; feature=Custom report builder; tier_required=Enterprise; quote=I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?
real_gapstranscript_id=TG-001; deal_alias=Deal-EC3025; amount=62000; feature=SCIM user provisioning; quote=We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.; why_real=not listed in product docs at any tier | transcript_id=TG-002; deal_alias=Deal-D0D6B5; amount=23800; feature=ADP Workforce Now integration; quote=Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.; why_real=supported HRIS list is Workday/BambooHR/Gusto/Rippling only
graded checksreal_gaps · plan_gate_classified · enablement_classified · gap_summary

Model's actual answer

show the model's full answer
Deal amounts and deal statuses are not provided. Active-deal status cannot be verified, and affected revenue cannot be calculated.

Candidate assessment

• Deal-EC3025 — Real gap. Amount: not provided.
  Prospect: “We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.”
  Evidence: TG-001; product_docs.md explicitly states SCIM user provisioning is not currently listed as supported.

• Deal-D0D6B5 — Real gap. Amount: not provided.
  Prospect: “Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.”
  Evidence: TG-002; product_docs.md explicitly states ADP Workforce Now is not currently listed as supported.

• Deal-CFE7F4 — Plan gate. Amount: not provided.
  Prospect: “I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?”
  Evidence: TG-003; product_docs.md lists the custom report builder as Enterprise-only. The prospect’s current plan and support for this specific report combination are not provided.

• Deal-84DBA6 — Rollout/enablement issue. Amount: not provided.
  Prospect: “We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.”
  Evidence: TG-004; product_docs.md lists Slack integration on all plans. The prospect describes missing training, not missing functionality.

• Deal-36C33F — Excluded: no prospect-raised gap. Amount: not provided.
  Prospect: “Good to know. The web version should be fine for our office staff for now.”
  Evidence: TG-005; the native-mobile-app claim comes from Alex Franklin, not the prospect. product_docs.md does not establish native mobile availability, so no gap classification is supported.

Real-gap summary

• SCIM user provisioning — Deal-EC3025: security approval blocker; amount missing.
• ADP Workforce Now integration — Deal-D0D6B5: prospect-described dealbreaker; amount missing.

Affected amount = amount for Deal-EC3025 + amount for Deal-D0D6B5 = unavailable because both amounts are missing.
rep-performance 5 tests
stale-pipeline-by-rep0.67
Plain-English question: Which open deals have had no email, call, or meeting in the last 7 days?
Prompt — what the model saw
The snapshot date is 2026-09-05. Using the open deals and the engagements-by-deal table, list every open deal with no logged email, call, or meeting in the last 7 days. Compute recency from the engagements table's last_email/last_call/last_meeting fields, NOT from any date field on the deal (the deal's last_contacted_field is unreliable). For each stale deal give: deal alias, owner name, stage, amount, days since last contact. Group by owner, order by amount descending within each owner. Then state, per owner, how many stale deals they carry and the total stale amount. (Do not write outreach text.)
cost $0.9352 · 111s · in 12 / out 6,628 tok
✓ stale_recall_ge_80✓ grouped_by_owner✓ jake_top_amount✓ kevin_stale_amt✗ jenn_stale_amt✗ stale_count

✓ Correct Answer (ground truth)

Plain-English answer: 65 open deals are stale; the largest is Deal-2D1F1B at $240,000, with Bryce Harmon carrying the most stale dollars at $626,243.
The answer in one line: jake_amt = 626,243; jenn_amt = 261,645; kevin_amt = 252,905.03; top_amount = 240,000; stale_count = 65; top_alias = Deal-2D1F1B
jake_amt626,243
jenn_amt261,645
kevin_amt252,905.03
stale_count65
top_amount240,000
top_aliasDeal-2D1F1B
graded checksstale_recall_ge_80 · grouped_by_owner · jake_top_amount · kevin_stale_amt · jenn_stale_amt · stale_count

Model's actual answer

show the model's full answer
As of 2026-09-05, stale means last logged contact was more than 7 days ago: before 2026-08-29. Last contact = latest nonblank last_email, last_call, or last_meeting on or before the snapshot date. Future-dated engagements are excluded; last_contacted_field is not used. Amounts are in the supplied units; currency was not specified.

Bryce Harmon

Deal alias  | Stage | Amount     | Days since last contact: snapshot − last contact
Deal-2D1F1B | DS1   | 240,000.00 | 2026-09-05 − 2026-06-16 = 81
Deal-66D1FC | DS1   |  99,000.00 | 2026-09-05 − 2026-08-20 = 16
Deal-950043 | DS1   |  70,000.00 | 2026-09-05 − 2026-08-17 = 19
Deal-B23205 | DS1   |  45,000.00 | 2026-09-05 − 2026-08-20 = 16
Deal-7BBDFA | DS3   |  37,440.00 | 2026-09-05 − 2026-07-21 = 46
Deal-332637 | DS2   |  36,000.00 | 2026-09-05 − 2026-08-27 = 9
Deal-1BEEBF | DS1   |  31,500.00 | 2026-09-05 − 2026-08-17 = 19
Deal-A414F6 | DS1   |  25,200.00 | 2026-09-05 − 2026-08-17 = 19
Deal-C5658B | DS1   |  23,400.00 | 2026-09-05 − 2026-08-20 = 16
Deal-40522D | DS3   |  21,000.00 | 2026-09-05 − 2026-08-17 = 19
Deal-C1FA6D | DS1   |  18,000.00 | 2026-09-05 − 2026-08-20 = 16
Deal-01E193 | DS1   |  12,600.00 | 2026-09-05 − 2026-08-28 = 8
Deal-F0EBBB | DS3   |  11,400.00 | 2026-09-05 − 2026-08-12 = 24
Deal-927338 | DS1   |  10,920.00 | 2026-09-05 − 2026-08-18 = 18
Deal-E25A09 | DS1   |   6,000.00 | 2026-09-05 − 2026-08-27 = 9
Deal-C9C286 | DS2   |   5,502.00 | 2026-09-05 − 2026-08-27 = 9
Deal-012CB1 | DS1   |       1.00 | 2026-09-05 − 2026-08-13 = 23
Deal-3795AD | DS2   |       1.00 | 2026-09-05 − 2026-08-28 = 8

Dana Mercer

Deal alias  | Stage | Amount    | Days since last contact: snapshot − last contact
Deal-44EA29 | DS2   | 60,000.00 | 2026-09-05 − 2026-08-26 = 10
Deal-E51FB7 | DS2   | 43,875.00 | 2026-09-05 − 2026-08-24 = 12
Deal-B42F46 | DS1   | 27,000.00 | 2026-09-05 − 2026-08-17 = 19
Deal-BA3DDC | DS3   | 23,400.00 | 2026-09-05 − 2026-08-21 = 15
Deal-9DDE86 | DS2   | 20,000.00 | 2026-09-05 − 2026-08-21 = 15
Deal-215CCA | DS3   | 18,900.00 | 2026-09-05 − 2026-08-19 = 17
Deal-5EED42 | DS3   | 16,250.00 | 2026-09-05 − 2026-08-25 = 11
Deal-57887A | DS2   | 15,000.00 | 2026-09-05 − 2026-08-28 = 8
Deal-944310 | DS4   | 10,500.00 | 2026-09-05 − 2026-08-03 = 33
Deal-3974EB | DS4   |  9,000.00 | 2026-09-05 − 2026-08-28 = 8
Deal-B7EBD1 | DS5   |  9,000.00 | 2026-09-05 − 2026-08-20 = 16
Deal-F40F04 | DS2   |  8,100.00 | 2026-09-05 − 2026-08-21 = 15
Deal-7599B8 | DS3   |  7,350.00 | 2026-09-05 − 2026-08-18 = 18
Deal-87DDD1 | DS1   |  5,000.00 | 2026-09-05 − 2026-08-17 = 19
Deal-F336B6 | DS3   |  4,200.00 | 2026-09-05 − 2026-08-21 = 15
Deal-0660B4 | DS4   |  1,920.00 | 2026-09-05 − 2026-08-20 = 16

Alex Franklin

Deal alias  | Stage | Amount    | Days since last contact: snapshot − last contact
Deal-CC08D1 | DS1   | 24,000.00 | 2026-09-05 − 2026-08-20 = 16
Deal-E73427 | DS3   | 18,000.00 | 2026-09-05 − 2026-08-26 = 10
Deal-885F45 | DS2   |  9,300.00 | 2026-09-05 − 2026-08-24 = 12
Deal-C2FF3C | DS1   |  8,316.00 | 2026-09-05 − 2026-08-26 = 10
Deal-0D2F7A | DS3   |  5,100.00 | 2026-09-05 − 2026-08-24 = 12
Deal-6C60D4 | DS3   |  4,800.00 | 2026-09-05 − 2026-08-24 = 12
Deal-13FEBD | DS2   |  4,680.00 | 2026-09-05 − 2026-08-24 = 12
Deal-819506 | DS1   |  4,400.00 | 2026-09-05 − 2026-08-28 = 8
Deal-9D0060 | DS3   |  3,840.00 | 2026-09-05 − 2026-08-24 = 12
Deal-690476 | DS2   |  3,600.00 | 2026-09-05 − 2026-08-18 = 18
Deal-C6D97A | DS4   |  3,240.00 | 2026-09-05 − 2026-08-28 = 8
Deal-EE195F | DS3   |  3,120.00 | 2026-09-05 − 2026-08-28 = 8
Deal-278DEC | DS3   |  2,700.00 | 2026-09-05 − 2026-08-28 = 8
Deal-635B8E | DS3   |  2,600.00 | 2026-09-05 − 2026-08-18 = 18
Deal-6883F3 | DS1   |  2,400.00 | 2026-09-05 − 2026-08-20 = 16
Deal-4A13AD | DS3   |  2,160.00 | 2026-09-05 − 2026-08-10 = 26
Deal-F67D31 | DS2   |  1,800.00 | 2026-09-05 − 2026-08-28 = 8
Deal-5FDCE4 | DS3   |  1,600.00 | 2026-09-05 − 2026-08-24 = 12
Deal-BA571A | DS4   |  1,080.00 | 2026-09-05 − 2026-08-18 = 18

Cole Ingram

Deal alias  | Stage | Amount    | Days since last contact: snapshot − last contact
Deal-D04904 | DS2   | 58,529.25 | 2026-09-05 − 2026-08-25 = 11
Deal-B25F40 | DS3   | 40,000.00 | 2026-09-05 − 2026-08-28 = 8
Deal-813836 | DS2   | 32,175.00 | 2026-09-05 − 2026-08-25 = 11
Deal-1BA595 | DS2   | 31,750.00 | 2026-09-05 − 2026-08-25 = 11
Deal-CFE1E8 | DS3   | 18,000.00 | 2026-09-05 − 2026-08-25 = 11
Deal-CD47A6 | DS2   | 12,168.00 | 2026-09-05 − 2026-08-25 = 11
Deal-627646 | DS3   | 11,193.00 | 2026-09-05 − 2026-08-25 = 11
Deal-FF809F | DS2   |  7,781.20 | 2026-09-05 − 2026-08-25 = 11
Deal-AF932D | DS2   |  7,225.40 | 2026-09-05 − 2026-08-25 = 11
Deal-A71728 | DS2   |  6,947.50 | 2026-09-05 − 2026-08-25 = 11
Deal-8BC9F5 | DS2   |  5,616.00 | 2026-09-05 − 2026-08-26 = 10
Deal-175395 | DS3   |  4,779.88 | 2026-09-05 − 2026-08-25 = 11
Deal-481E24 | DS3   |  4,140.00 | 2026-09-05 − 2026-08-26 = 10
Deal-C7F9BF | DS2   |  3,360.00 | 2026-09-05 − 2026-08-25 = 11
Deal-2F3A66 | DS3   |  3,334.80 | 2026-09-05 − 2026-08-25 = 11
Deal-342E96 | DS2   |  2,700.00 | 2026-09-05 − 2026-08-12 = 24
Deal-E568D5 | DS3   |  1,875.00 | 2026-09-05 − 2026-08-25 = 11
Deal-FD9F4E | DS5   |  1,330.00 | 2026-09-05 − 2026-08-26 = 10

Farid Osman

Deal alias  | Stage | Amount   | Days since last contact: snapshot − last contact
Deal-8BA24E | DS1   | 2,880.00 | 2026-09-05 − 2026-08-28 = 8
Deal-8FDCD2 | DS1   |     1.00 | 2026-09-05 − 2026-08-21 = 15

Missing engagement data — recency and stale status cannot be determined

These deals have no row in engagements_by_deal_90d.csv. They are not included in the confirmed-stale totals.

Owner         | Deal alias  | Stage | Amount   | Days since last contact
Alex Franklin | Deal-3EED2C | DS2   | 7,200.00 | Unknown
Elena Sinclair | Deal-57FF13 | DS1   | 2,100.00 | Unknown

Per-owner stale totals

Count = number of listed stale deals; total stale amount = sum of that owner's listed amounts.

Owner              | Confirmed stale deals | Total stale amount
Bryce Harmon         | 18                    | 692,964.00
Dana Mercer         | 16                    | 279,495.00
Alex Franklin      | 19                    | 106,736.00
Cole Ingram | 18                    | 252,905.03
Farid Osman        |  2                    |   2,881.00
Elena Sinclair      |  0 confirmed          |       0.00 confirmed

Amount arithmetic:
Bryce Harmon: 240,000 + 99,000 + 70,000 + 45,000 + 37,440 + 36,000 + 31,500 + 25,200 + 23,400 + 21,000 + 18,000 + 12,600 + 11,400 + 10,920 + 6,000 + 5,502 + 1 + 1 = 692,964.00
Dana Mercer: 60,000 + 43,875 + 27,000 + 23,400 + 20,000 + 18,900 + 16,250 + 15,000 + 10,500 + 9,000 + 9,000 + 8,100 + 7,350 + 5,000 + 4,200 + 1,920 = 279,495.00
Alex Franklin: 24,000 + 18,000 + 9,300 + 8,316 + 5,100 + 4,800 + 4,680 + 4,400 + 3,840 + 3,600 + 3,240 + 3,120 + 2,700 + 2,600 + 2,400 + 2,160 + 1,800 + 1,600 + 1,080 = 106,736.00
Cole Ingram: 58,529.25 + 40,000 + 32,175 + 31,750 + 18,000 + 12,168 + 11,193 + 7,781.20 + 7,225.40 + 6,947.50 + 5,616 + 4,779.88 + 4,140 + 3,360 + 3,334.80 + 2,700 + 1,875 + 1,330 = 252,905.03
Farid Osman: 2,880 + 1 = 2,881.00
activity-mix-vs-outcomePASS 1.00
Plain-English question: Per rep, what did the last 30 days of emails, calls, and meetings look like next to the deals that advanced?
Prompt — what the model saw
The snapshot date is 2026-09-05. For each sales rep (owner), compute the last-30-day totals of emails, calls, and meetings from the engagements table, and the number of deals that entered DS2 in the last 30 days (t_ds2 within 30 days of the snapshot). Report per rep: the activity mix as percentages (emails/calls/meetings share of that rep's total activities), total activities per DS2 entry, and rank the reps by that efficiency ratio (lowest activities per DS2 entry = most efficient). Name the most efficient rep and the highest-volume rep, and note if they differ. Do not attribute cause.
cost $0.8043 · 75s · in 12 / out 4,385 tok
✓ alex_ds2_30d✓ justin_volume_high✓ efficient_named✓ mix_present✓ no_causal_claim

✓ Correct Answer (ground truth)

Plain-English answer: Alex Franklin is the standout signal with 18 deals entering DS2 in 30 days (next closest: Bryce Harmon with 4), alongside per-rep activity totals.
The answer in one line: alex_ds2_30d = 18; bryce_ds2_30d = 4; note = legacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
alex_ds2_30d18
bryce_ds2_30d4
notelegacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
graded checksalex_ds2_30d · justin_volume_high · efficient_named · mix_present · no_causal_claim

Model's actual answer

show the model's full answer
Window: 2026-08-06 through 2026-09-05, inclusive (snapshot minus 30 days through snapshot). Activities use the supplied emails_30d, calls_30d, and meetings_30d columns, joined to owners by deal_id. DS2 entries count qualifying t_ds2 dates regardless of current stage.

Activity share = activity-type total ÷ total activities × 100.
Efficiency = total activities ÷ DS2 entries; lower ranks better.

| Rank | Rep | Emails | Calls | Meetings | Total activities: arithmetic | Activity mix: email / call / meeting | DS2 entries | Activities per DS2: arithmetic |
|---:|---|---:|---:|---:|---|---|---:|---|
| 1* | Alex Franklin | 307 | 36 | 41 | 307 + 36 + 41 = 384 | 79.95% / 9.38% / 10.68% | 18 | 384 ÷ 18 = 21.33 |
| 2 | Bryce Harmon | 162 | 0 | 43 | 162 + 0 + 43 = 205 | 79.02% / 0.00% / 20.98% | 4 | 205 ÷ 4 = 51.25 |
| 3 | Cole Ingram | 96 | 14 | 1 | 96 + 14 + 1 = 111 | 86.49% / 12.61% / 0.90% | 2 | 111 ÷ 2 = 55.50 |
| 4 | Farid Osman | 38 | 0 | 34 | 38 + 0 + 34 = 72 | 52.78% / 0.00% / 47.22% | 1 | 72 ÷ 1 = 72.00 |
| 5 | Dana Mercer | 84 | 18 | 11 | 84 + 18 + 11 = 113 | 74.34% / 15.93% / 9.73% | 1 | 113 ÷ 1 = 113.00 |
| — | Elena Sinclair | Missing | Missing | Missing | Unknown | N/A | 0 | Undefined |
| — | Gavin Porter (archived) | Missing | Missing | Missing | Unknown | N/A | 0 | Undefined |
| — | Hugo Lindqvist (archived) | Missing | Missing | Missing | Unknown | N/A | 0 | Undefined |

*Alex Franklin’s Deal-3EED2C entered DS2 on 2026-09-03 but has no engagements row. His activity totals and ranking therefore reflect recorded activity only; missing activity is not assumed to be zero. Elena Sinclair’s Deal-57FF13 also lacks an engagements row. Neither archived owner has deals or attributable engagements in the supplied data.

Most efficient on recorded activity: Alex Franklin, at 21.33 activities per DS2 entry.
Highest recorded activity volume: Alex Franklin, with 384 activities.
They do not differ.
ae-qtd-scorecardPASS 1.00
Plain-English question: Build Alex Franklin's QTD scorecard: bookings versus quota, pipeline, and conversion.
Prompt — what the model saw
Build a QTD scorecard for the named AE as of 2026-09-05: bookings vs quota with attainment percent, new vs expansion split, active pipeline by stage with amount, rolling 90-day DS2-to-won rate, win and loss counts with top loss reason, activity volume by type in the last 30 days. Exclude any closed-won deal dated before the quarter. Close with three coaching observations grounded in the numbers.
cost $0.6967 · 84s · in 12 / out 4,767 tok
✓ bookings✓ quota✓ attainment✓ new_expansion_split✓ top_loss_reason✓ coaching_obs

✓ Correct Answer (ground truth)

Plain-English answer: 75% attainment ($150,000 on a $200,000 quota), $113,500 new versus $36,500 expansion, pipeline by stage, and the rolling DS2-to-won rate.
The answer in one line: quota = 200000; bookings = 150000; new_bookings = 113500; expansion_bookings = 36500; open_deals_count = 125; attainment_pct = 75; owner = Alex Franklin (+12 more below)
quota200,000
attainment_pct75
bookings150,000
expansion_bookings36,500
expansion_count3
losses_qtd27
new_bookings113,500
new_count5
open_deals_count125
wins_qtd8
activity_30d_totalsemails_30d=807; calls_30d=112; meetings_30d=128; notes_30d=50
allowed_new18 items — e.g. Deal-A1C3E5, Deal-B7D2F4, Deal-C9E1A6, Deal-D4B8C2
ds2_to_wonwindow_start=2026-06-07; won_in_window=8; closed_in_window=35; rate=0.2286
ownerAlex Franklin
pipeline_by_stage{"DS1": {"count": 20, "amount": 284621}, "DS2": {"count": 28, "amount": 353760}, "DS3": {"count": 67, "amount": 552705}, "DS4": {"count": 5, "amount": 23574}, "DS5": {"count": 5, "amount": 45730}}
prior_quarter_excludeddeal_alias=Deal-B3E6F1; close_date=2026-06-20; amount=24000
quarter2026-Q3
snapshot_date2026-09-05
top_loss_reasonreason=Lost- Timing (1 year or more); count=13
graded checksbookings · quota · attainment · new_expansion_split · top_loss_reason · coaching_obs

Model's actual answer

show the model's full answer
Alex Franklin — QTD scorecard as of 2026-09-05
Quarter to date: 2026-07-01–2026-09-05.

Excluded throughout: Deal-B3E6F1, closed-won 2026-06-20 for $24,000, including its activity row.

Bookings vs quota
- New: Deal-A1C3E5 $40,000 + Deal-B7D2F4 $35,000 + Deal-C9E1A6 $21,000 + Deal-D4B8C2 $11,000 + Deal-E6F3A9 $6,500 = $113,500 across 5 wins.
- Expansion: Deal-F2C7D8 $20,000 + Deal-A8B4D6 $12,000 + Deal-C5D9E2 $4,500 = $36,500 across 3 wins.
- Total bookings: $113,500 + $36,500 = $150,000.
- Quota: $200,000; attainment = $150,000 ÷ $200,000 = 75.0%.
- Remaining quota: $200,000 − $150,000 = $50,000.
- Booking mix: new = $113,500 ÷ $150,000 = 75.7%; expansion = $36,500 ÷ $150,000 = 24.3%.

Active pipeline
All supplied open deals, including close dates beyond Q3; amounts are unweighted.

Stage    Deals       Amount
DS1         20     $284,621
DS2         28     $353,760
DS3         67     $552,705
DS4          5      $23,574
DS5          5      $45,730
Total      125   $1,260,390

Amount reconciliation: $284,621 + $353,760 + $552,705 + $23,574 + $45,730 = $1,260,390.
Of this, 22 open deals totaling $109,363 have Q3 close dates.

Rolling 90-day DS2-to-won rate
Window: 2026-06-08–2026-09-05, inclusive.
Using deals that entered DS2 in this window:
- 8 won + 27 lost + 76 still open = 111 deals.
- DS2-to-won conversion as of the snapshot = 8 ÷ 111 = 7.2%.
- For comparison, excluding unresolved deals gives a closed-outcome win rate of 8 ÷ (8 + 27) = 22.9%; this is not the full-cohort conversion rate.

QTD wins and losses
- Wins: 8.
- Losses: 27.
- Top loss reason by count: “Lost- Timing (1 year or more)” — 13 ÷ 27 = 48.1% of losses, totaling $184,681.

Activity volume — last 30 days
Using the supplied _30d fields as the snapshot’s 2026-08-07–2026-09-05 window. Individual activity timestamps were not provided.

Type        Lost deals + Open deals + QTD wins = Total
Emails             109 + 599 + 89 = 797
Calls               25 +  54 + 31 = 110
Meetings            13 +  90 + 23 = 126
Notes               25 +   1 + 21 = 47

Total: 797 + 110 + 126 + 47 = 1,080 recorded activities.

Three coaching observations
1. Focus on the $50,000 quota gap: Q3-dated open pipeline is $109,363 ÷ $50,000 = 2.19× the gap, not guaranteed bookings. Revalidate close dates and next steps, starting with Deal-7A2454: $1,275 remains open despite a 2026-09-04 close date.
2. Tighten timing qualification: 13 of 27 losses (48.1%), representing $184,681, cite “Lost- Timing (1 year or more).” Confirm a buyer-owned implementation date before advancing opportunities.
3. Validate late-stage buyer engagement: Deal-547B2B ($11,200) and Deal-A2B47C ($6,360) are DS5 with 2026-09-11 close dates, but each has zero recorded calls and meetings in the last 30 days. Secure live next steps for this $11,200 + $6,360 = $17,560 of pipeline.
multithreading-gapPASS 1.00
Plain-English question: Which open deals are single-threaded or under-threaded?
Prompt — what the model saw
Find every open deal that is single-threaded (fewer than 2 active contacts) or under-threaded (fewer than 3, or all contacts in one persona). Active means engaged in the last 60 days and not marked former. Personas: economic buyer, champion, HR admin, IT security, finance. For each flagged deal: amount, stage, active contact count, personas present, personas missing, the most valuable persona to add given the stage, and any on-file unengaged contact at that company who fits; otherwise say none on file.
cost $0.4222 · 55s · in 9 / out 2,478 tok
✓ flagged_recall✓ persona_language✓ single_thread_term✓ on_file_fit

✓ Correct Answer (ground truth)

Plain-English answer: Deals with fewer than 2 active contacts (single-threaded) or fewer than 3, or all in one persona (under-threaded), each flagged with the missing persona.
The answer in one line: active_window_start = 2026-07-07
active_window_start2026-07-07
allowed_new59 items — e.g. C-077A0E, C-0D15DF, C-1BDFE0, C-2AE3AA
best_persona_ruleDS2/DS3: economic buyer > IT security > finance > HR admin > champion; DS4/DS5: economic buyer > finance > IT security > HR admin > champion; first persona not present among active contacts
flag_ruleflag if active_contacts < 2 (single-threaded) or active_contacts < 3 or all active contacts share one persona (under-threaded); active = engaged within 60 days and not former
flaggeddeal_alias=Deal-36C33F; amount=15000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-50D386; amount=36000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5408B0; amount=14850; stage=DS4; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5BFE3B; amount=18000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-885F45; amount=9300; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=IT security | deal_alias=Deal-92D97D; amount=60000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer … (+5 more)
healthy_not_flaggedDeal-4B0BEB, Deal-84DBA6, Deal-D348E1
persona_universeeconomic buyer, champion, HR admin, IT security, finance
graded checksflagged_recall · persona_language · single_thread_term · on_file_fit

Model's actual answer

show the model's full answer
Assuming an as-of date of 2026-09-13: 2026-09-13 − 60 days = 2026-07-15. Active contacts must have engaged on or after that cutoff and not be former.

Missing for every deal: amount, stage, and open/closed status. Therefore, these are threading flags among the supplied deals—not a verified list of open deals. The most valuable persona to add given stage cannot be determined for any deal. Contacts below fit missing personas; they are not stage-ranked recommendations.

Flagged: 5 single-threaded + 6 additional under-threaded = 11 deals. Single-threaded deals also meet the under-threaded definition.

Active-count arithmetic: listed contacts − former contacts − stale non-former contacts.

1. Deal-EC3025 — C-FDD0C7
   Flag: single-threaded.
   Active: 2 − 1 − 0 = 1.
   Present: champion.
   Missing: economic buyer, HR admin, IT security, finance.
   On-file unengaged fit: CT-6827DB — Chief People Officer, economic buyer.

2. Deal-92D97D — C-E23238
   Flag: single-threaded.
   Active: 2 − 0 − 1 = 1.
   Present: HR admin.
   Missing: economic buyer, champion, IT security, finance.
   On-file fit: CT-A902AE — Head of Employee Experience, champion; dormant contact in deal_contacts.csv, last engaged 2026-06-01. None in unengaged_contacts.csv.

3. Deal-50D386 — C-EB10E4
   Flag: under-threaded; fewer than 3 active contacts.
   Active: 2 − 0 − 0 = 2.
   Present: champion, HR admin.
   Missing: economic buyer, IT security, finance.
   On-file unengaged fit: CT-A1C4B3 — Chief People Officer, economic buyer.

4. Deal-D0D6B5 — C-32918E
   Flag: under-threaded; all active contacts in one persona.
   Active: 3 − 0 − 0 = 3.
   Present: champion.
   Missing: economic buyer, HR admin, IT security, finance.
   On-file unengaged fit: CT-1FA4DB — Chief People Officer, economic buyer.

5. Deal-5BFE3B — C-535D36
   Flag: under-threaded; fewer than 3 active contacts and one persona.
   Active: 2 − 0 − 0 = 2.
   Present: champion.
   Missing: economic buyer, HR admin, IT security, finance.
   On-file unengaged fit: none on file.

6. Deal-36C33F — C-077A0E
   Flag: single-threaded.
   Active: 3 − 2 − 0 = 1.
   Present: IT security.
   Missing: economic buyer, champion, HR admin, finance.
   On-file unengaged fit: CT-1DB73E — Chief People Officer, economic buyer. Former contacts excluded.

7. Deal-885F45 — C-5E8EFB
   Flag: under-threaded; fewer than 3 active contacts.
   Active: 2 − 0 − 0 = 2.
   Present: economic buyer, champion.
   Missing: HR admin, IT security, finance.
   On-file unengaged fit: CT-B3F25D — IT Security Lead, IT security.

8. Deal-FCBE5B — C-737030
   Flag: single-threaded.
   Active: 1 − 0 − 0 = 1.
   Present: champion.
   Missing: economic buyer, HR admin, IT security, finance.
   On-file unengaged fit: none on file.

9. Deal-5408B0 — C-2AE3AA
   Flag: under-threaded; fewer than 3 active contacts.
   Active: 2 − 0 − 0 = 2.
   Present: champion, HR admin.
   Missing: economic buyer, IT security, finance.
   On-file unengaged fit: CT-07FA76 — Chief People Officer, economic buyer.

10. Deal-C6D97A — C-5A8FC2
    Flag: under-threaded; all active contacts in one persona.
    Active: 3 − 0 − 0 = 3.
    Present: champion.
    Missing: economic buyer, HR admin, IT security, finance.
    On-file unengaged fit: none on file.

11. Deal-F9A08A — C-0D15DF
    Flag: single-threaded.
    Active: 2 − 0 − 1 = 1.
    Present: champion.
    Missing: economic buyer, HR admin, IT security, finance.
    On-file unengaged fit: CT-697541 — Chief People Officer, economic buyer. Also available for re-engagement: CT-913581 — Chief People Officer, economic buyer; last engaged 2026-06-20.
talk-track-from-transcriptsPASS 1.00
Plain-English question: From this rep's last 10 calls: how do they open, handle objections, and how often do they lock a next step?
Prompt — what the model saw
From this rep's last 10 calls: what they lead with in the first five minutes, how they handle the three most common objections, how often a concrete next step is agreed (rate), and every competitor a prospect raised. Quote at most one line per finding and only lines that appear in the transcripts. End with two coaching notes.
cost $0.3876 · 38s · in 9 / out 1,543 tok
✓ objection_count✓ next_step_rate✓ competitors✓ coaching✓ quote_used

✓ Correct Answer (ground truth)

Plain-English answer: 8 of 10 calls open with the same customer-win story; objection handling, next-step agreement rate, and every competitor raised, with quotes.
The answer in one line: n_transcripts = 10; objection_count = 3; rep = Alex Franklin (+3 more below)
n_transcripts10
objection_count3
allowed_new13 items — e.g. TT-001, TT-002, TT-003, TT-004
competitor_quotesAwardco=We're also in late talks with Awardco — their rewards catalog looks bigger than yours.; Kudos=How are you different from Kudos? Our CEO used them at her last company.
competitors_prospect_raisedAwardco, Kudos
next_stepagreed=7; total=10; rate=0.7
objectionsobjection=budget; count=4; quote=Honestly, budget is locked until next fiscal year — I can't add a new line item right now.; transcript_id=TT-001 | objection=timing; count=3; quote=This is good, but can we revisit it next quarter? Open enrollment starts in October and we're underwater.; transcript_id=TT-002 | objection=in-house; count=3; quote=We already do recognition with a spreadsheet and quarterly gift cards — why would we change?; transcript_id=TT-004
opener_themetheme=opens with a customer recognition win story (400-person retailer, 18% turnover reduction); count=8; total=10; quote=Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it.; transcript_id=TT-001
repAlex Franklin
rep_raised_competitors_excludedWorkhuman
graded checksobjection_count · next_step_rate · competitors · coaching · quote_used

Model's actual answer

show the model's full answer
First five minutes

Alex Franklin leads with a customer-outcome story in 8 ÷ 10 = 80% of calls: a retailer’s claimed 18% reduction in regretted turnover after automating milestone awards. These are rep-stated results, not independently verified data.
- Calls: Deal-D348E1, Deal-5408B0, Deal-547B2B, Deal-C61CF7, Deal-60C2C2, Deal-EDC141, Deal-D9A12F, Deal-84DBA6; all at minute 0.
- Exceptions: Deal-403845 opens with security review then pricing; Deal-1E2498 opens with pricing at the prospect’s prior request. Each is 1 ÷ 10 = 10%.
- Early competitive positioning: he introduces Workhuman at minute 2 in Deal-C61CF7. In Deal-547B2B and Deal-EDC141, prospects introduce competitors at minute 4; he responds at minute 5.

Three most common objections and his handling

1. Budget/approval constraints: 5 ÷ 10 = 50% of calls when committee blockers are included. Explicit locked-budget objections occur in 4 ÷ 10 = 40%: Deal-D348E1, Deal-547B2B, Deal-60C2C2, Deal-84DBA6. Deal-403845 adds a budget-committee blocker; Deal-84DBA6 also mentions committee approval but is counted only once.
   He counters locked budgets with turnover savings and the retailer’s claimed avoided-backfill savings:
   “Totally fair. Most teams fund this out of turnover savings — that retailer saved about $210k in avoided backfills, which is how their finance team signed off.” — Deal-D348E1, minute 8.
   When committee approval blocks progress in Deal-403845 and Deal-84DBA6, he acknowledges it without securing a follow-up.

2. Timing/workload: 3 ÷ 10 = 30% — Deal-5408B0, Deal-C61CF7, Deal-D9A12F.
   He acknowledges open-enrollment workload and proposes a smaller pilot to generate internal evidence:
   “Makes sense. What if we scope a 90-day pilot with one department so you have internal data before next quarter's planning?” — Deal-5408B0, minute 8.
   All three agree to a working session; the transcripts do not establish agreement to run the pilot.

3. Status quo/no reason to change: 3 ÷ 10 = 30% — Deal-403845, Deal-EDC141, Deal-1E2498.
   He contrasts spreadsheets and gift cards with automation and recognition analytics:
   “Spreadsheets work until they scale — the difference is automation: milestones fire without HR lifting a finger, and you get analytics on who is being recognized.” — Deal-403845, minute 8.
   He does not visibly explore the prospect’s current administrative burden or quantify its impact. Deal-EDC141 later states there is no urgency; he accepts that without further discovery.

Concrete next-step agreement rate

Counting an explicit prospect agreement to a specific activity and time: 7 ÷ 10 × 100 = 70%.

- Agreed: Deal-D348E1, Deal-5408B0, Deal-547B2B, Deal-C61CF7, Deal-60C2C2, Deal-D9A12F, Deal-1E2498. Each agrees to a working session Thursday at 2pm and to bring the HRIS manager.
- Evidence: “Yes, Thursday at 2pm works — send the invite and I'll bring our HRIS manager.” — Deal-D348E1, minute 15.
- No concrete next step agreed: Deal-403845, Deal-EDC141, Deal-84DBA6.
- The transcripts do not confirm that invitations were sent or sessions occurred.

Every competitor raised by a prospect

- Awardco — Deal-547B2B, minute 4: “We're also in late talks with Awardco — their rewards catalog looks bigger than yours.”
- Kudos — Deal-EDC141, minute 4: “How are you different from Kudos? Our CEO used them at her last company.”

Workhuman is excluded: Alex Franklin, not the prospect, introduces it in Deal-C61CF7.

Two coaching notes

1. Follow the outcome story with prospect-specific discovery. Before repeating savings or automation claims, establish the prospect’s turnover costs, manual recognition workload, and reason to change—especially in Deal-403845 and Deal-EDC141.
2. Replace passive acceptance of committee or urgency blockers with a mutually agreed follow-up tied to the decision process: identify the decision owner, required evidence, and review date. Apply this to Deal-403845, Deal-EDC141, and Deal-84DBA6.
reporting-analytics 5 tests
q3-weighted-forecastPASS 1.00
Plain-English question: Build the Q3 2026 weighted forecast from the open pipeline, and show the math deal by deal.
Prompt — what the model saw
Below are the open deals in the default sales pipeline with a close date between 2026-07-01 and 2026-10-15, pulled from the CRM warehouse on 2026-09-05. Stages: DS1 (discovery) through DS5 (contract). Build the Q3 2026 forecast (quarter = 2026-07-01 to 2026-09-30). Weighting rule: 100 percent of COMMIT plus 35 percent of BEST_CASE. PIPELINE deals count zero. Only deals with a close date inside the quarter count. Report: COMMIT total, BEST_CASE total, weighted forecast, count of deals in each category inside the quarter, the deals you excluded for being outside the quarter (count and total amount), and the top 5 BEST_CASE deals by amount inside the quarter. Show arithmetic. Then, in a final '## Data quality' section of no more than four sentences, name the data quality problems in this extract that would make the forecast unreliable if it were run unattended.
cost $0.5164 · 60s · in 9 / out 3,497 tok
✓ commit_total_44729✓ best_case_total_203565✓ weighted_115976.75✓ commit_count_7✓ best_case_count_24✓ excluded_count_32✓ excluded_amount_227575✓ top5_aliases✓ flags_blank_owner✓ flags_missing_why_buys

✓ Correct Answer (ground truth)

Plain-English answer: A stage-weighted Q3 forecast number computed from the open pipeline, with per-stage math that reconciles to the deal data.
The answer in one line: excluded_amount = 227575; best_case_total = 203565; weighted_forecast = 115,976.75; commit_total = 44729; excluded_out_of_quarter = 32; best_case_count = 24; flags = blank owner column + empty why_buys fields must be surfaced (+2 more below)
excluded_amount227,575
best_case_count24
best_case_total203,565
commit_count7
commit_total44,729
excluded_out_of_quarter32
weighted_forecast115,976.75
flagsblank owner column + empty why_buys fields must be surfaced
top5_best_caseDeal-2D7423, Deal-25F752, Deal-E53952, Deal-5EED42, Deal-FA32A0
graded checkscommit_total_44729 · best_case_total_203565 · weighted_115976.75 · commit_count_7 · best_case_count_24 · excluded_count_32 · excluded_amount_227575 · top5_aliases · flags_blank_owner · flags_missing_why_buys

Model's actual answer

show the model's full answer
## Q3 2026 forecast

Source: provided deals.csv; close dates included from 2026-07-01 through 2026-09-30. Amounts are in the extract’s unspecified currency.

| Category | Q3 deal count | Q3 total | Weight | Forecast contribution |
|---|---:|---:|---:|---:|
| COMMIT | 7 | 44,729.00 | 100% | 44,729.00 |
| BEST_CASE | 24 | 203,565.00 | 35% | 71,247.75 |
| PIPELINE | 23 | 201,637.40 | 0% | 0.00 |

COMMIT arithmetic:
11,200 + 9,000 + 9,000 + 6,360 + 5,400 + 2,520 + 1,249 = 44,729.

BEST_CASE arithmetic:
38,935 + 24,000 + 19,656 + 16,250 + 11,116 + 10,800 + 10,500 + 9,890 + 9,720 + 9,000 + 7,200 + 3,840 + 3,780 + 3,600 + 3,240 + 3,150 + 3,120 + 3,060 + 2,916 + 2,760 + 2,484 + 2,100 + 1,920 + 528 = 203,565.

Weighted forecast = (100% × 44,729) + (35% × 203,565) + (0% × 201,637.40)
= 44,729 + 71,247.75 + 0
= **115,976.75**.

Included deal count = 7 + 24 + 23 = 54.

## Excluded: outside Q3

32 deals totaling 227,575.00, all with October close dates.

Count arithmetic: 1 COMMIT + 9 BEST_CASE + 22 PIPELINE = 32.
Amount arithmetic: 13,770 COMMIT + 28,240 BEST_CASE + 185,565 PIPELINE = 227,575.

| Deal | Close date | Amount |
|---|---|---:|
| Deal-E51FB7 | 2026-10-01 | 43,875 |
| Deal-B936FE | 2026-10-09 | 18,000 |
| Deal-D9A12F | 2026-10-15 | 17,000 |
| Deal-D348E1 | 2026-10-15 | 13,770 |
| Deal-4062CF | 2026-10-15 | 10,800 |
| Deal-293AF3 | 2026-10-09 | 9,000 |
| Deal-034D49 | 2026-10-15 | 9,000 |
| Deal-E0ADD8 | 2026-10-15 | 7,920 |
| Deal-9F2E43 | 2026-10-08 | 7,690 |
| Deal-FCBE5B | 2026-10-07 | 7,500 |
| Deal-712010 | 2026-10-15 | 7,200 |
| Deal-6691E0 | 2026-10-15 | 5,700 |
| Deal-C61CF7 | 2026-10-09 | 5,400 |
| Deal-600CD9 | 2026-10-02 | 5,400 |
| Deal-A92065 | 2026-10-15 | 5,400 |
| Deal-1D532E | 2026-10-15 | 5,400 |
| Deal-48B656 | 2026-10-15 | 5,160 |
| Deal-E531A6 | 2026-10-15 | 4,800 |
| Deal-D1E6C2 | 2026-10-09 | 4,400 |
| Deal-D9E112 | 2026-10-09 | 4,300 |
| Deal-5AD94B | 2026-10-15 | 4,000 |
| Deal-901332 | 2026-10-15 | 3,600 |
| Deal-47AE31 | 2026-10-09 | 3,600 |
| Deal-15D24F | 2026-10-09 | 3,600 |
| Deal-766C74 | 2026-10-14 | 3,300 |
| Deal-ED725A | 2026-10-08 | 2,400 |
| Deal-8AD4A5 | 2026-10-07 | 1,800 |
| Deal-D7E999 | 2026-10-15 | 1,800 |
| Deal-ED13B0 | 2026-10-09 | 1,680 |
| Deal-5FDCE4 | 2026-10-01 | 1,600 |
| Deal-7FA0C3 | 2026-10-01 | 1,400 |
| Deal-F5A622 | 2026-10-08 | 1,080 |

## Top 5 BEST_CASE deals inside Q3

| Rank | Deal | Amount | Close date |
|---|---|---:|---|
| 1 | Deal-2D7423 | 38,935 | 2026-09-30 |
| 2 | Deal-25F752 | 24,000 | 2026-09-25 |
| 3 | Deal-E53952 | 19,656 | 2026-09-30 |
| 4 | Deal-5EED42 | 16,250 | 2026-09-30 |
| 5 | Deal-FA32A0 | 11,116 | 2026-09-25 |

## Data quality

Owner is blank on 85 of 86 deals; only Deal-C9C286 has an owner, limiting accountability and validation.  
71 deals have why_buys_chars = 0, including all five largest Q3 BEST_CASE deals, leaving no recorded why-buys support.  
Deal-A5E80A is DS1/COMMIT and Deal-499BF6 is DS2/COMMIT—early-stage commitments requiring validation, although both remain included under the stated category rule.  
Deal-333EBB (2026-08-28), Deal-57FF13 (2026-09-02), Deal-31AD2C (2026-09-04), and Deal-7A2454 (2026-09-04) remain open with close dates preceding the 2026-09-05 extraction date, indicating stale close dates or statuses.
aha-moment-2x2PASS 1.00
Plain-English question: Do companies that hit 5+ givers and one redemption in month one actually retain better at 24 months?
Prompt — what the model saw
Test the activation hypothesis on the company cohort below: companies with 5+ unique givers (m1_users >= 5) AND 1+ successful redemption (m1_redemptions >= 1) in their first calendar month retain better at 24 months than those with one signal or neither. Every company in this extract is already 25+ months old, so current_status = 'active' means retained at 24 months. Report the full 2x2 (cohort size and 24-month retention rate for: both signals, givers-only, redemption-only, neither), state how many companies were excluded from the denominator and why (if none, say so), name the single signal with the largest retention lift, and state what this does and does not prove.
cost $0.6771 · 43s · in 18 / out 1,203 tok
✓ both_66✓ givers_47✓ redemp_31✓ neither_40✓ excluded_0✓ largest_lift_givers

✓ Correct Answer (ground truth)

Plain-English answer: The four cohort cells are graded against computed 24-month retention; the numbers come from the scorer's cohort table, not opinion.
The answer in one line: both_5plus_givers_and_redemption = 66.0% (31 of 47)
both_5plus_givers_and_redemption66.0% (31 of 47)
exclusionsnone — claim bounded as correlation, not causation
givers_only46.9% (23 of 49)
largest_single_signal_liftunique givers
neither40.0% (38 of 95)
redemption_only31.0% (9 of 29)
graded checksboth_66 · givers_47 · redemp_31 · neither_40 · excluded_0 · largest_lift_givers

Model's actual answer

show the model's full answer
Using current_status = 'active' as retained at 24 months:

| First-calendar-month signals | Cohort size | Retained | 24-month retention |
|---|---:|---:|---:|
| Both: m1_users ≥ 5 AND m1_redemptions ≥ 1 | 47 | 31 | 31 ÷ 47 = 65.96% |
| Givers-only: m1_users ≥ 5 AND m1_redemptions = 0 | 49 | 23 | 23 ÷ 49 = 46.94% |
| Redemption-only: m1_users < 5 AND m1_redemptions ≥ 1 | 29 | 9 | 9 ÷ 29 = 31.03% |
| Neither: m1_users < 5 AND m1_redemptions = 0 | 95 | 38 | 38 ÷ 95 = 40.00% |

Denominator: 47 + 49 + 29 + 95 = 220 companies. Excluded: 0. All are 25+ months old and have the required signal and status fields. Cancelled and non_renewing companies remain in the denominator but are not counted as retained; missing industry/country values do not affect eligibility.

Largest single-signal lift: 5+ unique givers. Comparing each single-signal-only group with neither:
- Givers-only: 46.94% − 40.00% = +6.94 percentage points.
- Redemption-only: 31.03% − 40.00% = −8.97 percentage points.

Conclusion: The hypothesis is supported descriptively in this extract. Both signals have the highest retention: +19.02 percentage points versus givers-only, +34.92 versus redemption-only, and +25.96 versus neither.

This demonstrates an observed association, not causation. It does not prove that driving companies across these thresholds will improve retention, that these thresholds are optimal, or that the result generalizes beyond this cohort. Confounding has not been controlled for, and statistical significance has not been established.
arr-rollforward-reconciliationPASS 1.00
Plain-English question: Does CRM ARR match billing ARR — and if not, exactly where is the variance?
Prompt — what the model saw
Reconcile ARR between CRM company records and billing subscriptions as of 2026-09-05. Billing ARR is MRR x 12. Report both totals, the variance, and a decomposition into buckets (status mismatch, rounding, missing records, other) that sums exactly to the variance. List each mismatched account with a suggested owner. Business rule: any subscription with a term other than 12 months must have cf_agreement_end_date populated; list violations.
cost $0.4415 · 48s · in 9 / out 2,755 tok
✓ billing_total✓ crm_total✓ variance✓ buckets_sum✓ violations_listed

✓ Correct Answer (ground truth)

Plain-English answer: Billing $604,739.28 vs CRM $603,581.76: a $1,157.52 variance decomposed into status-mismatch, rounding, and missing-record buckets.
The answer in one line: n_companies = 40; n_mismatched_accounts = 7; as_of = 2026-09-05 (+4 more below)
n_companies40
n_mismatched_accounts7
as_of2026-09-05
billing_arr_total604739.28
bucket_sum_equals_variance1
buckets{"status_mismatch": {"amount": "-13158.48", "accounts": [{"company_alias": "C-0C8323BF", "billing_arr": "0.00", "hubspot_arr": "4905.24", "diff": "-4905.24", "suggested_owner": "Dana Mercer"}, {"compa
crm_arr_total603581.76
ruleterm_months != 12 requires cf_agreement_end_date
variance1157.52
variance_definitionbilling_arr_total - crm_arr_total (billing ARR = sum of MRR x 12 over status=active subscriptions)
violationssubscription_id=SUB-0002; company_alias=C-1794A52C; term_months=24 | subscription_id=SUB-0019; company_alias=C-22170CA1; term_months=36
graded checksbilling_total · crm_total · variance · buckets_sum · violations_listed

Model's actual answer

show the model's full answer
ARR reconciliation — 2026-09-05

Basis: Treat supplied statuses as the as-of snapshot; effective-date history was not provided. Active billing ARR = MRR × 12; cancelled subscriptions contribute zero. Currency is unspecified. Variance = CRM − billing.

Totals
- CRM ARR: sum of hubspot_arr = 603,581.76
- Billing ARR: active MRR 50,394.94 × 12 = 604,739.28
- Variance: 603,581.76 − 604,739.28 = −1,157.52

Variance decomposition

| Bucket | Arithmetic | Variance |
|---|---|---:|
| Status mismatch | 4,905.24 + 8,253.24 | +13,158.48 |
| Rounding (provisional) | 16.00 + 20.00 | +36.00 |
| Missing records | 16,497.24 − 28,449.24 | −11,952.00 |
| Other | 24,396.00 − 26,796.00 | −2,400.00 |
| Total | 13,158.48 + 36.00 − 11,952.00 − 2,400.00 | −1,157.52 |

Rounding is a provisional classification: both CRM values match billing ARR rounded to the nearest hundred, but no rounding policy was provided.

Mismatched accounts

| Company alias | CRM ARR | Billing ARR arithmetic | CRM − billing | Issue | Suggested owner |
|---|---:|---|---:|---|---|
| C-0C8323BF | 4,905.24 | Cancelled → 0.00 | +4,905.24 | CRM retains ARR for cancelled subscription | RevOps + Billing Ops |
| C-0DC4FB8C | 8,253.24 | Cancelled → 0.00 | +8,253.24 | CRM retains ARR for cancelled subscription | RevOps + Billing Ops |
| C-0D66DF9E | 23,200.00 | 1,932.00 × 12 = 23,184.00 | +16.00 | Apparent rounding; verify policy | RevOps |
| C-14D70CE0 | 18,200.00 | 1,515.00 × 12 = 18,180.00 | +20.00 | Apparent rounding; verify policy | RevOps |
| C-0D5BBE3A | 16,497.24 | No billing record supplied | +16,497.24 | CRM-only record | Billing Ops + RevOps |
| C-21629AA4 | Missing | 2,370.77 × 12 = 28,449.24 | −28,449.24 | Billing-only record | RevOps |
| C-0F7269D7 | 24,396.00 | 2,233.00 × 12 = 26,796.00 | −2,400.00 | Unexplained ARR difference | Finance + RevOps |

Missing records contribute zero on the absent side solely for reconciliation; this does not establish that actual ARR is zero. Owners are suggested roles, not provided assignments.

Agreement-end-date violations

| Subscription | Company alias | Term months | Violation |
|---|---|---:|---|
| SUB-0002 | C-1794A52C | 24 | cf_agreement_end_date is blank |
| SUB-0019 | C-22170CA1 | 36 | cf_agreement_end_date is blank |
kvm-month-over-monthPASS 1.00
Plain-English question: How did the four core KVMs move in August versus July?
Prompt — what the model saw
Report the four core KVMs for 2026-08 against 2026-07: giving rate, redemptions per user, 1:1 meetings engagement, pulse check engagement. For each: value, prior value, absolute and relative change, direction. Name the KVM with the largest relative move and the segment (plan_tier or size_band) driving it if the data supports that; otherwise say the data does not. Under 200 words plus one table.
cost $0.3532 · 33s · in 6 / out 1,894 tok
✓ four_kvms✓ biggest_mover_named✓ direction_correct✓ segment_named✓ word_cap

✓ Correct Answer (ground truth)

Plain-English answer: Giving rate 60.27% (up 0.07%), redemptions per user 1.7302 (up 0.01%), with 1:1 meetings and pulse engagement each up fractionally — direction reported for all four.
The answer in one line: n_company_rows_per_month = 30; aggregation = simple mean of company-level values per month
n_company_rows_per_month30
aggregationsimple mean of company-level values per month
biggest_moverpulse_engagement
driving_segmentsize_band=enterprise; 2026-07=0.55; 2026-08=0.2743; rel_change_pct=-50.13
kvms{"giving_rate": {"2026-07": 0.6023, "2026-08": 0.6027, "abs_change": 0.0004, "rel_change_pct": 0.07, "direction": "up"}, "redemptions_per_user": {"2026-07": 1.73, "2026-08": 1.7302, "abs_change": 0.00
segment_detail{"smb": {"2026-07": 0.6588, "2026-08": 0.6573, "rel_change_pct": -0.23}, "mid_market": {"2026-07": 0.593, "2026-08": 0.5942, "rel_change_pct": 0.2}, "enterprise": {"2026-07": 0.55, "2026-08": 0.2743,
graded checksfour_kvms · biggest_mover_named · direction_correct · segment_named · word_cap

Model's actual answer

show the model's full answer
Values are unweighted means across the same 30 companies; user counts are missing, so user-weighted KVMs cannot be calculated. Source: kvm_monthly.csv.

| KVM | 2026-08 | 2026-07 | Absolute change | Relative change | Direction |
|---|---:|---:|---:|---:|---|
| Giving rate | 60.2713% | 60.2297% | +0.0417 pp | +0.0692% | Up |
| Redemptions per user | 1.730163 | 1.729983 | +0.000180 | +0.0104% | Up |
| 1:1 meetings engagement | 44.7177% | 44.6887% | +0.0290 pp | +0.0649% | Up |
| Pulse check engagement | 50.8610% | 60.0587% | −9.1977 pp | −15.3145% | Down |

Arithmetic: mean = monthly sum ÷ 30; absolute change = August − July; relative change = absolute change ÷ July × 100. Percentage-point changes multiply raw-rate differences by 100.

August/July sums:
- Giving: 18.0814 / 18.0689.
- Redemptions: 51.9049 / 51.8995.
- 1:1: 13.4153 / 13.4066.
- Pulse: 15.2583 / 18.0176.

Pulse check engagement had the largest relative move. The enterprise size_band drives the decline: 54.998% → 27.428%, or −27.570 pp (−50.129%). Its contribution is (10 ÷ 30) × −27.570 = −9.190 pp, accounting for 99.917% of the net decline. All companies are tier_three, so plan_tier cannot distinguish the driver.
rewards-ytd-sectionPASS 1.00
Plain-English question: Produce the redemption section of the weekly rewards report through August.
Prompt — what the model saw
Produce the redemption section of the weekly rewards report through the last completed month (2026-08): redemption count, spend, unique redeemers, redemptions per redeemer, provider mix as percent of spend, top 5 countries by redemptions. Say which month is the last completed month. Provider shares must sum to 100.
cost $1.0767 · 71s · in 15 / out 3,854 tok
✓ count✓ spend✓ unique_redeemers✓ month_named✓ provider_mix✓ top_country

✓ Correct Answer (ground truth)

Plain-English answer: 378 redemptions, $27,846 spend, 235 unique redeemers at 1.61 each, custom rewards at 39% of spend, US leading all countries.
The answer in one line: redemption_count = 378; unique_redeemers = 235; redemptions_per_redeemer = 1.61; total_spend_usd = 27846.00
redemption_count378
redemptions_per_redeemer1.61
unique_redeemers235
provider_mix_pct_of_spendTangoCard=11.6; Tremendous=30.54; custom=39.05; Snappy=18.81
top5_countries_by_countcountry=US; count=244 | country=CA; count=24 | country=AU; count=21 | country=GB; count=17 | country=NL; count=17
total_spend_usd27846.00
graded checkscount · spend · unique_redeemers · month_named · provider_mix · top_country

Model's actual answer

show the model's full answer
Redemptions — YTD through August 2026

Last completed month: 2026-08 (August). Reporting period: January 1–August 31, 2026. Source: supplied redemptions_ytd.csv.

- Redemption count: 90 + 192 + 37 + 59 = 378
- Spend: $3,230 + $8,505 + $10,873 + $5,238 = $27,846 USD
- Unique redeemers: 235 distinct user_key values across all providers
- Redemptions per redeemer: 378 ÷ 235 = 1.61

Provider mix (% of spend):
- TangoCard: $3,230 ÷ $27,846 × 100 = 11.60%
- Tremendous: $8,505 ÷ $27,846 × 100 = 30.54%
- custom: $10,873 ÷ $27,846 × 100 = 39.05%
- Snappy: $5,238 ÷ $27,846 × 100 = 18.81%
- Total: 11.60% + 30.54% + 39.05% + 18.81% = 100.00%

Top 5 countries by redemption count:
1. US: 244
2. CA: 24
3. AU: 21
4. GB: 17 (tied)
4. NL: 17 (tied)
customer-success 5 tests
churn-save-eligibilityPASS 1.00
Plain-English question: Which at-risk accounts actually qualify for a churn-save offer, and which play fits each?
Prompt — what the model saw
Which at-risk accounts qualify for a churn-save offer under the documented eligibility rules, what amount is at stake per account and in total, and which play fits each (usage revival, executive touch, commercial concession)? Cite the signal that justifies each play. List accounts that look at risk but do not qualify and why.
cost $0.3157 · 38s · in 6 / out 1,679 tok
✓ eligible_set✓ total_at_stake✓ plays_cited✓ noneligible_named✓ rules_applied

✓ Correct Answer (ground truth)

Plain-English answer: Only accounts meeting all three rules qualify — e.g. C-0F6C0F34 at $49,707 (executive touch) and C-0B827671 at $25,365 (usage revival) — with amount and play per account.
The answer in one line: n_accounts = 30; snapshot_date = 2026-09-05
n_accounts30
eligibleaccount_alias=C-0F6C0F34; amount_at_stake=49707.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0B827671; amount_at_stake=25365.00; play=usage revival; justifying_signal=usage_trend_3m=declining over the last 3 months | account_alias=C-0B360C78; amount_at_stake=35748.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 75% - risk is commercial, not adoption | account_alias=C-0B0F1BAB; amount_at_stake=5494.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0CA21961; amount_at_stake=16829.00; play=usage revival; justifying_signal=seat utilization 26% is below 50% | account_alias=C-0E9C27D1; amount_at_stake=41235.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 85% - risk is commercial, not adoption … (+2 more)
non_eligible_at_riskaccount_alias=C-0BC71BDD; health_score=55 | account_alias=C-0BA71F12; health_score=52 | account_alias=C-0F6694C3; health_score=43 | account_alias=C-0BE96399; health_score=54 | account_alias=C-0F876796; health_score=47 | account_alias=C-0FCCD2DF; health_score=43 … (+1 more)
ruleshealth_score < 60, churn_save_eligible_amount > 0, renewal within 120 days of snapshot
snapshot_date2026-09-05
total_amount_at_stake224601.00
graded checkseligible_set · total_at_stake · plays_cited · noneligible_named · rules_applied

Model's actual answer

show the model's full answer
8 accounts qualify, with $224,601.00 in churn-save eligible amount at stake—not total ARR.

Eligibility requires all three rules: health_score < 60 (R1), churn_save_eligible_amount > $0 (R2), and renewal within 120 days of 2026-09-05, through 2027-01-03 (R3). Days below = renewal_date − 2026-09-05.

### Qualifying accounts

| Account | Health | Renewal (days away) | Eligible amount at stake | Recommended play and supporting signal |
|---|---:|---|---:|---|
| C-0F6C0F34 | 51 | 2026-10-03 (28) | $49,707.00 | Executive touch: champion_active=false despite growing usage. |
| C-0B827671 | 56 | 2026-11-14 (70) | $25,365.00 | Usage revival: declining usage; seats used = 113 ÷ 202 = 55.9%. |
| C-0B360C78 | 57 | 2026-10-28 (53) | $35,748.00 | Undetermined: usage is growing and champion is active. Low health establishes risk, but no commercial issue is provided to justify a concession. |
| C-0B0F1BAB | 38 | 2026-09-23 (18) | $5,494.00 | Executive touch: champion_active=false, with flat usage and an imminent renewal. |
| C-0CA21961 | 58 | 2026-12-28 (114) | $16,829.00 | Usage revival: flat usage and seats used = 84 ÷ 325 = 25.8%. |
| C-0E9C27D1 | 39 | 2026-09-24 (19) | $41,235.00 | Undetermined: usage is flat, seats used = 134 ÷ 157 = 85.4%, and champion is active. No commercial issue is provided to justify a concession. |
| C-0CEF69FD | 53 | 2026-11-21 (77) | $32,621.00 | Executive touch: champion_active=false despite growing usage. |
| C-0D3278C7 | 54 | 2026-11-12 (68) | $17,602.00 | Usage revival: declining usage; seats used = 126 ÷ 380 = 33.2%. |

Total arithmetic:
$49,707 + $25,365 + $35,748 + $5,494 + $16,829 + $41,235 + $32,621 + $17,602 = $224,601.00.

Play assignments are recommendations, not documented eligibility rules. No pricing objection, budget constraint, or other commercial signal is supplied for any account; commercial concession cannot be justified from these data alone. Eligible amounts are not documented concession sizes.

### At-risk accounts that do not qualify

All below meet R1 but fail R2 and/or R3.

| Account | Health | Eligible-amount field | Renewal (days away) | Disqualifying rule(s) |
|---|---:|---:|---|---|
| C-0BC71BDD | 55 | $0.00 | 2026-10-27 (52) | R2: amount is zero. |
| C-0BA71F12 | 52 | $6,824.00 | 2027-04-11 (218) | R3: renewal exceeds 120 days. |
| C-0F6694C3 | 43 | $0.00 | 2027-03-21 (197) | R2: amount is zero; R3: renewal exceeds 120 days. |
| C-0BE96399 | 54 | $0.00 | 2026-10-29 (54) | R2: amount is zero. |
| C-0F876796 | 47 | $19,958.00 | 2027-02-06 (154) | R3: renewal exceeds 120 days. |
| C-0FCCD2DF | 43 | $0.00 | 2027-04-23 (230) | R2: amount is zero; R3: renewal exceeds 120 days. |
| C-10A56B0F | 54 | $0.00 | 2026-12-12 (98) | R2: amount is zero. |
expansion-kitPASS 1.00
Plain-English question: Build the expansion kit for the customer who just replied to our outreach.
Prompt — what the model saw
A customer replied to outreach (reply included). Build the expansion kit: seat coverage (licensed over headcount), two lines on usage health, headroom in seats and ARR at the current per-seat rate, who replied and whether they can buy, the right buyer from our contacts, and a reply email under 150 words that moves toward the right person without being pushy and cites one real usage fact.
cost $0.2411 · 19s · in 6 / out 573 tok
✓ coverage_pct✓ headroom_arr✓ right_buyer✓ replier_not_buyer✓ usage_fact✓ email_under_150

✓ Correct Answer (ground truth)

Plain-English answer: 37.5% seat coverage, 250 headroom seats worth $15,000 ARR at $60 a seat, usage up 43.2% in six months — and the replier cannot buy; Dana R. (VP People) is the right buyer.
The answer in one line: headroom_seats = 250; coverage_pct = 37.5; company_alias = C-0DDFC9A7 (+3 more below)
headroom_seats250
coverage_pct37.5
company_aliasC-0DDFC9A7
headroom_arr15000.00
per_seat_annual_rate60.00
replierMaria S.
replier_can_buy0
right_buyername=Dana R.; title=VP People
usage_factsactive_users_2026_08=126; active_users_2026_03=88; growth_pct_6m=43.2; licensed_seat_utilization_pct=84
usage_health_linesMonthly active users grew from 88 (2026-03) to 126 (2026-08), +43.2% over 6 months., Last month 126 of 150 licensed seats were active (84.0% utilization), approaching the seat cap.
graded checkscoverage_pct · headroom_arr · right_buyer · replier_not_buyer · usage_fact · email_under_150

Model's actual answer

show the model's full answer
C-0DDFC9A7

Seat coverage: 150 licensed / 400 headcount × 100 = 37.5%.

Usage health:
- Monthly active users increased every month from March to August: 88 → 126; growth = (126 − 88) / 88 × 100 = 43.2%.
- August utilization: 126 / 150 × 100 = 84%; 150 − 126 = 24 licensed seats not active that month.

Expansion headroom:
- Seats: 400 − 150 = 250 additional seats.
- Current annual per-seat rate: $9,000 / 150 = $60.
- Potential incremental ARR: 250 × $60 = $15,000, assuming full-headcount licensing at the current rate—not confirmed demand.

Who replied: Maria S., People Operations Coordinator, on 2026-09-02. She explicitly cannot make purchasing decisions and offered an introduction.

Right buyer: Dana R., VP People; Maria S. confirms she owns budget and seat expansion. Last engaged: 2026-05-18.

Reply email

Subject: RE: Growing your team's recognition program

Hi Maria,

Thanks for the context—glad the team is enjoying Bonusly. Monthly active users grew from 88 in March to 126 in August, so there’s a useful usage story to share with Dana R.

If you’re still comfortable making an introduction, I’d appreciate it. I can send Dana a brief usage summary and current seat coverage to help her assess whether any changes would be useful. No need to decide on expansion now.

Thanks for your help,
Cole
onboarding-statusPASS 1.00
Plain-English question: What is done and what is missing in this customer's onboarding before the mid-onboarding call?
Prompt — what the model saw
Prep the CSM for the mid-onboarding call: what is complete (integrations connected, allowance set, admins added, first recognition given, first redemption), what is not, early engagement signals, and the three things to cover on the call. Mark nothing complete without a data field that shows it.
cost $0.2376 · 20s · in 6 / out 540 tok
✓ complete_items✓ incomplete_items✓ three_items✓ grounded

✓ Correct Answer (ground truth)

Plain-English answer: 25 days live: Slack connected, allowance set, admins added — but the HRIS integration is not connected, with early engagement signals and three call priorities.
The answer in one line: days_live = 25; company_alias = C-0D284E42
days_live25
call_itemsConnect the HRIS integration (integration_hris is empty) to automate user provisioning and offboarding., Get to first redemption (first_redemption_at is empty): walk admins through the rewards catalog and confirm allowance funding., Build the giving habit: daily active givers averaged 4.3 in week 1 vs 13.0 last week - agree on manager nudges and a recognition cadence.
checklist{"slack_integration_connected": {"complete": true, "field": "integration_slack", "value": "2026-08-12"}, "hris_integration_connected": {"complete": false, "field": "integration_hris", "value": ""}, "a
company_aliasC-0D284E42
engagement_signalsavg_active_givers_week1=4.3; avg_active_givers_last_week=13; latest_active_givers=15; days_with_activity=25
graded checkscomplete_items · incomplete_items · three_items · grounded

Model's actual answer

show the model's full answer
C-0D284E42 — Mid-onboarding call prep

Complete, supported by data:
- Slack connected: integration_slack = 2026-08-12.
- Allowance set: allowance_set = 2026-08-13.
- Admins added: admins_added = 2.
- First recognition given: first_recognition_at = 2026-08-15 14:22.

Not confirmed complete:
- HRIS connected: integration_hris is blank.
- First redemption: first_redemption_at is blank.
Blank fields indicate missing evidence, not proof these actions never occurred.

Early engagement signals:
- Daily active givers increased from 3 on 2026-08-11 to 15 on 2026-09-04: 15 − 3 = 12 more; 12 ÷ 3 × 100 = 400% growth.
- First seven days averaged 30 ÷ 7 = 4.29 active givers/day; last seven days averaged 91 ÷ 7 = 13.
- Employee count and unique-giver totals are missing, so adoption percentage and distinct participants cannot be determined.
- Data discrepancy to clarify: active givers appear before the recorded first recognition on 2026-08-15.

Three things to cover on the call:
1. Confirm HRIS connection status; identify any blocker and agree on an owner and next step.
2. Confirm whether a redemption has occurred; if not, walk through the redemption process and identify barriers.
3. Review rising giver activity and how to sustain it; clarify the active_givers definition versus the first-recognition timestamp and obtain employee count to assess adoption.
renewal-risk-conflicting-datesPASS 1.00
Plain-English question: Two systems disagree on renewal dates — which do you trust per account, and who is actually renewing inside 90 days?
Prompt — what the model saw
Produce a 90-day renewal risk brief. Two systems hold renewal dates and disagree on some accounts; decide which to trust per account and say why (multi-year contracts are known to be wrong in ChurnZero). For every renewal: company, CSM, ARR, date used, seat utilization, 3-month usage trend, risk rating with one sentence of evidence. Flag every disagreement. Close with total ARR renewing and ARR at risk.
cost $0.5677 · 68s · in 9 / out 4,067 tok
✓ total_renewing✓ arr_at_risk✓ disagreements_flagged✓ trust_rule

✓ Correct Answer (ground truth)

Plain-English answer: Chargebee wins per account (multi-year contracts are known wrong in ChurnZero), with each trusted date, the reason, and a 90-day-window flag.
The answer in one line: n_accounts = 20; n_disagreements = 5; snapshot_date = 2026-09-05 (+2 more below)
n_accounts20
n_disagreements5
accounts20 items — e.g. account_alias=C-0B144C78; csm=Cole Ingram; arr=30899.00; trusted_renewal_date=2026-11-02; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.4; usage_3m_ratio=1.03; risk=low; evidence=3-month usage ratio 1.03 (last3 avg 103 vs prior3 100), seat utilization 75% | account_alias=C-0B20DB64; csm=Dana Mercer; arr=21770.00; trusted_renewal_date=2026-10-07; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=56.6; usage_3m_ratio=1; risk=medium; evidence=3-month usage ratio 1.00 (last3 avg 295 vs prior3 295), seat utilization 57% | account_alias=C-0B344485; csm=Elena Sinclair; arr=64384.00; trusted_renewal_date=2026-11-16; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=78; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 241 vs prior3 231), seat utilization 78% | account_alias=C-0B369871; csm=Dana Mercer; arr=85128.00; trusted_renewal_date=2026-10-29; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.1; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 330 vs prior3 316), seat utilization 75%
arr_at_risk_90d359409.00
disagreementsaccount_alias=C-0B7D2C30; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-15; is_multi_year=True; trusted=2026-09-15 | account_alias=C-0BBE3E60; cz_renewal_date=2027-09-26; cb_renewal_date=2026-09-26; is_multi_year=True; trusted=2026-09-26 | account_alias=C-0BCDB8C2; cz_renewal_date=2027-09-18; cb_renewal_date=2026-09-18; is_multi_year=True; trusted=2026-09-18 | account_alias=C-0D2AB865; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-22; is_multi_year=True; trusted=2026-09-22 | account_alias=C-0F5D2323; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-29; is_multi_year=True; trusted=2026-09-29
snapshot_date2026-09-05
total_arr_renewing_90d1048715.00
trust_rulemulti-year contracts: Chargebee is authoritative (ChurnZero known wrong); otherwise systems agree or Chargebee wins
window2026-09-05 to 2026-12-04
graded checkstotal_renewing · arr_at_risk · disagreements_flagged · trust_rule

Model's actual answer

show the model's full answer
90-day renewal risk brief

Scope: Assumed window September 1–November 29, 2026, using the latest usage month, August 2026, as the reporting cutoff; no as-of date was supplied. All 20 reconciled renewals fall within this window. ARR currency was not supplied.

Date reconciliation

Trust Chargebee for the five multi-year accounts below because ChurnZero is known to misstate multi-year renewal dates. Every disagreement is flagged:

| Company | ChurnZero date rejected | Chargebee date used | Reason |
|---|---|---|---|
| C-0B7D2C30 | 2026-09-10 | 2026-09-15 | 36-month multi-year contract |
| C-0BCDB8C2 | 2027-09-18 | 2026-09-18 | 36-month multi-year contract |
| C-0D2AB865 | 2026-09-10 | 2026-09-22 | 24-month multi-year contract |
| C-0BBE3E60 | 2027-09-26 | 2026-09-26 | 24-month multi-year contract |
| C-0F5D2323 | 2026-09-10 | 2026-09-29 | 24-month multi-year contract |

For every other account, use the shared date: Chargebee and ChurnZero agree, and the contract is 12 months. “Agreed” below identifies these accounts individually.

Calculation and rating rules

- Seat utilization = ChurnZero seats_used ÷ seats × 100.
- Three-month trend shows June → July → August active users; percentage change = (August − June) ÷ June × 100. Example: C-0B7D2C30 = (84 − 97) ÷ 97 × 100 = −13.4%.
- Analyst-defined risk rules, not supplied company policy: High if utilization is below 40% or usage declines at least 10%; otherwise Medium if utilization is below 70%; otherwise Low.
- ARR at risk includes High + Medium, at full ARR—not probability-weighted expected loss.

Renewal detail

| Company | CSM | ARR | Date used / basis | Seat utilization | Jun → Jul → Aug; change | Risk — one-sentence evidence |
|---|---|---:|---|---|---|---|
| C-0B7D2C30 | Dana Mercer | 65,901 | 2026-09-15 / CB override | 274/476 = 57.6% | 97 → 94 → 84; −13.4% | High — Usage fell 13.4% while only 57.6% of seats were utilized. |
| C-0BCDB8C2 | Cole Ingram | 54,427 | 2026-09-18 / CB override | 232/424 = 54.7% | 127 → 118 → 110; −13.4% | High — Usage fell 13.4% with 54.7% seat utilization. |
| C-0D2AB865 | Elena Sinclair | 38,022 | 2026-09-22 / CB override | 250/407 = 61.4% | 125 → 117 → 109; −12.8% | High — Usage declined each month, falling 12.8% overall. |
| C-0BBE3E60 | Dana Mercer | 30,993 | 2026-09-26 / CB override | 74/114 = 64.9% | 39 → 35 → 33; −15.4% | High — Usage fell 15.4% despite 64.9% seat utilization. |
| C-0F5D2323 | Cole Ingram | 90,647 | 2026-09-29 / CB override | 111/390 = 28.5% | 20 → 21 → 18; −10.0% | High — Seat utilization is only 28.5%, and usage fell 10.0%. |
| C-0EC6999D | Elena Sinclair | 79,419 | 2026-10-03 / Agreed | 31/112 = 27.7% | 17 → 16 → 15; −11.8% | High — Seat utilization is only 27.7%, and usage declined 11.8%. |
| C-0B20DB64 | Dana Mercer | 21,770 | 2026-10-07 / Agreed | 214/378 = 56.6% | 294 → 298 → 294; 0.0% | Medium — Usage was flat overall, but seat utilization is only 56.6%. |
| C-0BBC4E7A | Cole Ingram | 56,374 | 2026-10-10 / Agreed | 228/337 = 67.7% | 142 → 141 → 139; −2.1% | Medium — Seat utilization is 67.7%, with a modest 2.1% usage decline. |
| C-0FD551AB | Elena Sinclair | 48,815 | 2026-10-14 / Agreed | 210/376 = 55.9% | 123 → 122 → 126; +2.4% | Medium — Usage grew 2.4%, but seat utilization remains 55.9%. |
| C-0F9F8F13 | Dana Mercer | 46,230 | 2026-10-18 / Agreed | 199/352 = 56.5% | 185 → 185 → 182; −1.6% | Medium — Seat utilization is 56.5%, accompanied by a slight usage decline. |
| C-0BC34584 | Cole Ingram | 16,740 | 2026-10-22 / Agreed | 327/494 = 66.2% | 104 → 104 → 106; +1.9% | Medium — Usage increased 1.9%, but seat utilization is below 70%. |
| C-0B7A7546 | Elena Sinclair | 35,062 | 2026-10-25 / Agreed | 182/205 = 88.8% | 64 → 65 → 63; −1.6% | Low — Seat utilization is 88.8%, with only a 1.6% usage decline. |
| C-0B369871 | Dana Mercer | 85,128 | 2026-10-29 / Agreed | 317/422 = 75.1% | 326 → 330 → 333; +2.1% | Low — Seat utilization is 75.1%, and usage rose each month. |
| C-0B144C78 | Cole Ingram | 30,899 | 2026-11-02 / Agreed | 169/224 = 75.4% | 101 → 101 → 106; +5.0% | Low — Seat utilization is 75.4%, with 5.0% usage growth. |
| C-0FC4DBB8 | Elena Sinclair | 94,732 | 2026-11-05 / Agreed | 356/464 = 76.7% | 189 → 191 → 193; +2.1% | Low — Seat utilization is 76.7%, and usage increased each month. |
| C-0D5BBE3A | Dana Mercer | 39,740 | 2026-11-09 / Agreed | 85/102 = 83.3% | 88 → 90 → 91; +3.4% | Low — Seat utilization is 83.3%, with 3.4% usage growth. |
| C-0FB9D5AF | Cole Ingram | 63,158 | 2026-11-13 / Agreed | 144/199 = 72.4% | 173 → 173 → 176; +1.7% | Low — Seat utilization exceeds 70%, and usage increased 1.7%. |
| C-0B344485 | Elena Sinclair | 64,384 | 2026-11-16 / Agreed | 224/287 = 78.0% | 238 → 240 → 244; +2.5% | Low — Seat utilization is 78.0%, and usage grew each month. |
| C-0CB2C1B4 | Dana Mercer | 40,628 | 2026-11-20 / Agreed | 386/473 = 81.6% | 47 → 48 → 49; +4.3% | Low — Reported seat utilization is 81.6% and usage grew 4.3%, although August active users were only 49 versus 386 seats_used. |
| C-22170CA1 | Cole Ingram | 45,646 | 2026-11-24 / Agreed | 251/294 = 85.4% | 143 → 148 → 146; +2.1% | Low — Seat utilization is 85.4%, and August usage remained 2.1% above June. |

Data limitations: seats_used has no snapshot date or definition linking it to monthly active users. August active users exceed seats_used for C-0B20DB64, C-0B369871, C-0D5BBE3A, C-0FB9D5AF, and C-0B344485; these measures cannot safely be treated as interchangeable. Ratings are usage-based screens, not confirmed renewal outcomes; customer intent and commercial health were not supplied.

Sources: churnzero_renewals.csv for CSM, ARR and seats; chargebee_terms.csv for contract terms and reconciled dates; usage_12m.csv for active-user trends.

ARR totals

- High, 6 accounts: 65,901 + 54,427 + 38,022 + 30,993 + 90,647 + 79,419 = 359,409.
- Medium, 5 accounts: 21,770 + 56,374 + 48,815 + 46,230 + 16,740 = 189,929.
- Low, 9 accounts: 35,062 + 85,128 + 30,899 + 94,732 + 39,740 + 63,158 + 64,384 + 40,628 + 45,646 = 499,377.

Total ARR renewing = 359,409 + 189,929 + 499,377 = 1,048,715.
ARR at risk = 359,409 + 189,929 = 549,338, or 549,338 ÷ 1,048,715 × 100 = 52.4%.
ticket-theme-synthesis0.80
Plain-English question: What are the real themes in this quarter's support tickets, ignoring the existing tags?
Prompt — what the model saw
Synthesize the quarter's support tickets into 4 to 7 themes. Do not trust existing tags; read the text. Per theme: name, count, share, distinct accounts, ARR affected, two ticket ids, one-line recommendation. Separate broad patterns from single-account noise and rank by ARR exposure, not volume.
cost $0.4129 · 39s · in 9 / out 1,733 tok
✗ theme_count_range✓ counts_verifiable✓ single_account_flagged✓ arr_ranked✓ ticket_ids

✓ Correct Answer (ground truth)

Plain-English answer: 80 tickets in 4-7 themes — points-not-posting leads (20 tickets, 25%, $31,100 ARR affected) — each with counts, distinct accounts, ARR, and a recommendation.
The answer in one line: total_tickets = 80; quarter_window = 2026-06-01 through 2026-08-31
total_tickets80
noteexisting tags are noisy/unreliable; classify by text
quarter_window2026-06-01 through 2026-08-31
ranking_by_arr_exposurehris_provisioning, redemption_failures, billing_dispute, points_not_posting, slack_sync
sample_ticket_ids80 items — e.g. IC-460004, IC-460016, IC-460001, IC-460041
single_account_noise_themebilling_dispute
themes{"points_not_posting": {"count": 20, "share_pct": 25.0, "distinct_accounts": 9, "arr_affected": "31100.00", "example_ticket_ids": ["IC-460001", "IC-460002"], "single_account": false}, "slack_sync": {"
graded checkstheme_count_range · counts_verifiable · single_account_flagged · arr_ranked · ticket_ids

Model's actual answer

show the model's full answer
Scope: supplied June–August 2026 tickets. Themes classified from body_text, ignoring tags. Share = theme tickets ÷ 80 × 100. ARR affected counts each distinct account once per theme; it is exposure, not confirmed lost revenue. Ranked by ARR, not ticket volume.

1. HRIS provisioning failures — $114,000 ARR
   Count: 12 | Share: 12 ÷ 80 × 100 = 15.0% | Distinct accounts: 3
   ARR: $36,000 + $30,000 + $48,000 = $114,000.
   Pattern: Cross-account failures to create new-hire accounts, including skipped hires without logged errors. C-0B2213A9 contributes 7 of 12 tickets, but the issue also affects C-0F6C0F34 and C-0DDFC9A7.
   Tickets: IC-460059, IC-460060.
   Recommendation: Reconcile HRIS new hires against created accounts, replay missed provisioning, and alert on silent skips.

2. Redemption and gift-card failures — $68,800 ARR
   Count: 18 | Share: 18 ÷ 80 × 100 = 22.5% | Distinct accounts: 7
   ARR: $8,900 + $10,700 + $9,600 + $8,700 + $11,000 + $9,600 + $10,300 = $68,800.
   Pattern: Broad failures spanning checkout hangs, missing gift-card delivery, and points deducted despite failed orders.
   Tickets: IC-460025, IC-460024.
   Recommendation: Trace checkout through fulfillment and reconcile deducted points against unsuccessful orders.

3. Invoice seat-count and tier-price errors — $52,000 ARR
   Count: 16 | Share: 16 ÷ 80 × 100 = 20.0% | Distinct accounts: 1
   ARR: C-0E9C27D1 alone = $52,000, counted once despite 16 tickets.
   Pattern: Single-account concentration, not evidence of a broad billing defect; repeated seat-count discrepancies and incorrect renewal-tier pricing warrant account-specific escalation.
   Tickets: IC-460069, IC-460078.
   Recommendation: Assign one billing owner to reconcile licensed seats and contracted tier pricing, correct invoices, and verify the next billing cycle.

4. Recognition points not credited — $31,100 ARR
   Count: 20 | Share: 20 ÷ 80 × 100 = 25.0% | Distinct accounts: 9
   ARR: $3,500 + $4,500 + $4,500 + $2,700 + $3,400 + $4,200 + $2,900 + $2,500 + $2,900 = $31,100.
   Pattern: Widest account reach: delivered recognitions without points, stale balances, and team-wide posting failures.
   Tickets: IC-460004, IC-460016.
   Recommendation: Reconcile delivered recognitions to points-ledger entries, repair missing credits, and monitor posting delays.

5. Slack integration failures — $18,900 ARR
   Count: 14 | Share: 14 ÷ 80 × 100 = 17.5% | Distinct accounts: 4
   ARR: $4,400 + $5,400 + $3,900 + $5,200 = $18,900.
   Pattern: Cross-account recognition-sync, persistent-authentication, and slash-command failures; C-0BA71F12 contributes 6 of 14 tickets, but three other accounts also report failures.
   Tickets: IC-460047, IC-460051.
   Recommendation: Investigate authentication and sync-setting persistence, then test recognition delivery and slash commands across affected accounts.

Reconciliation: 12 + 18 + 16 + 20 + 14 = 80 tickets; 3 + 7 + 1 + 9 + 4 = 24 distinct accounts, with no cross-theme account overlap; total affected ARR = $114,000 + $68,800 + $52,000 + $31,100 + $18,900 = $284,800.
marketing 5 tests
case-study-matchPASS 1.00
Plain-English question: Which three existing customers are the best case-study matches for this prospect?
Prompt — what the model saw
Given the prospect, rank the three most similar existing customers with public case studies for use as social proof. Similarity on industry, size band, use case, region. Cite the fields that drove each match. Do not name any customer without a case study and never include billing data or contact names.
cost $0.2488 · 23s · in 6 / out 600 tok
✓ top3✓ decoys_excluded✓ cites_fields✓ no_billing

✓ Correct Answer (ground truth)

Plain-English answer: Three ranked matches on industry, size, use case, and region (e.g. C-11C31562), with the higher-scoring decoy explicitly excluded for a mismatched field.
The answer in one line: prospect_alias = C-82AF3719
constraintsno_billing_data, no_contact_names, only_customers_with_has_case_study_true_named
decoys_excludedcompany_alias=C-D6217CAA; similarity_score=4; reason=no_case_study | company_alias=C-C153868F; similarity_score=4; reason=no_case_study
prospect_aliasC-82AF3719
prospect_fields_usedindustry, size_band, use_case, region
top3company_alias=C-11C31562; similarity_score=3 | company_alias=C-64171065; similarity_score=3 | company_alias=C-A13C193D; similarity_score=3
graded checkstop3 · decoys_excluded · cites_fields · no_billing

Model's actual answer

show the model's full answer
For prospect C-82AF3719, the three closest customers tie for first using equally weighted exact matches on industry, size_band, use_case, and region (1 point each). All have has_case_study=true in customers.csv.

1. C-64171065 — tied first, 3/4 = 1+1+1+0.
   Matches: industry=Technology; size_band=Mid-Market; use_case=employee_recognition.
   Difference: region=NA-East versus prospect’s NA-West.

1. C-A13C193D — tied first, 3/4 = 1+1+0+1.
   Matches: industry=Technology; size_band=Mid-Market; region=NA-West.
   Difference: use_case=retention versus prospect’s employee_recognition.

1. C-11C31562 — tied first, 3/4 = 0+1+1+1.
   Matches: size_band=Mid-Market; use_case=employee_recognition; region=NA-West.
   Difference: industry=Manufacturing versus prospect’s Technology.

Case-study URLs and explicit confirmation of public availability were not provided; eligibility relies on has_case_study=true.
channel-attribution-mathPASS 1.00
Plain-English question: Compute cost per SQM, cost per SQO, and pipeline per dollar for every channel over six months.
Prompt — what the model saw
Compute channel performance for the trailing 6 months. Paid: spend, SQMs, SQOs, cost per SQM, cost per SQO, SQM-to-SQO rate, pipeline amount, pipeline per dollar. Organic: volume, SQO rate, pipeline. Report a channel with spend and zero SQMs as undefined, not zero. Flag rows where SQO date precedes SQM date. Finish with a reallocation recommendation and your confidence given sample sizes.
cost $0.5746 · 78s · in 12 / out 4,270 tok
✓ undefined_channel✓ undefined_not_zero✓ flagged_rows✓ best_channel✓ reallocation

✓ Correct Answer (ground truth)

Plain-English answer: LinkedIn Ads at $960/SQM and $3,000/SQO, organic search at $90,000 pipeline on zero spend, paid social flagged at $18,000 with zero SQMs, and a channel winner named.
The answer in one line: best_pipeline_per_dollar_channel = paid_search
best_pipeline_per_dollar_channelpaid_search
confidence_notewebinars and referral have small SQM samples (<20); treat rates as directional.
flagged_rows_sqo_before_sqmCT-000041, CT-000044
months2026-03, 2026-04, 2026-05, 2026-06, 2026-07, 2026-08
per_channel{"linkedin_ads": {"spend_usd": 24000, "sqms": 25, "sqos": 8, "pipeline_amount": 96000, "cost_per_sqm": 960.0, "cost_per_sqo": 3000.0, "sqm_to_sqo_rate": 0.32, "pipeline_per_dollar": 4.0}, "organic_sea
reallocation_recommendationfrom=paid_social; to=paid_search; rationale=paid_social has spend but zero SQMs (cost per SQM undefined, not zero); paid_search has the highest pipeline per dollar.
undefined_channelspaid_social
graded checksundefined_channel · undefined_not_zero · flagged_rows · best_channel · reallocation

Model's actual answer

show the model's full answer
Period: March–August 2026, the six months provided. Volume means contacts with an SQM date in this period; SQOs have a populated SQO date in the period. Metrics below retain flagged records as reported.

Paid performance

| Channel | Spend | SQMs | SQOs | Cost/SQM | Cost/SQO | SQM→SQO | Pipeline | Pipeline/$ |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| paid_search | $36,000 | 40 | 18 | $900 | $2,000 | 45.00% | $720,000 | 20.00x |
| linkedin_ads | $24,000 | 25 | 8 | $960 | $3,000 | 32.00% | $96,000 | 4.00x |
| paid_social | $18,000 | 0 | 0 | Undefined | Undefined | Undefined | $0 | 0.00x |
| webinars | $9,000 | 12 | 5 | $750 | $1,800 | 41.67% | $60,000 | 6.67x |

Arithmetic:
- paid_search: spend = 6 × $6,000 = $36,000; cost/SQM = $36,000 ÷ 40 = $900; cost/SQO = $36,000 ÷ 18 = $2,000; conversion = 18 ÷ 40 = 45%; pipeline = 18 × $40,000 = $720,000; pipeline/$ = $720,000 ÷ $36,000 = 20x.
- linkedin_ads: spend = 6 × $4,000 = $24,000; cost/SQM = $24,000 ÷ 25 = $960; cost/SQO = $24,000 ÷ 8 = $3,000; conversion = 8 ÷ 25 = 32%; pipeline = 8 × $12,000 = $96,000; pipeline/$ = $96,000 ÷ $24,000 = 4x.
- paid_social: spend = 6 × $3,000 = $18,000; no contact rows. Cost/SQM and cost/SQO = $18,000 ÷ 0, undefined; conversion = 0 ÷ 0, undefined. Recorded pipeline/$ = $0 ÷ $18,000 = 0x, not an undefined denominator.
- webinars: spend = 6 × $1,500 = $9,000; cost/SQM = $9,000 ÷ 12 = $750; cost/SQO = $9,000 ÷ 5 = $1,800; conversion = 5 ÷ 12 = 41.67%; pipeline = 5 × $12,000 = $60,000; pipeline/$ = $60,000 ÷ $9,000 = 6.67x.

Organic performance

| Channel | Volume (SQMs) | SQOs | SQO rate | Pipeline |
|---|---:|---:|---:|---:|
| organic_search | 30 | 10 | 10 ÷ 30 = 33.33% | 10 × $9,000 = $90,000 |
| referral | 15 | 6 | 6 ÷ 15 = 40.00% | 6 × $8,000 = $48,000 |

Organic costs are not provided; these channels cannot be assumed cost-free.

Date-quality flags

- CT-000044, linkedin_ads: SQO 2026-07-18 precedes SQM 2026-07-23.
- CT-000041, linkedin_ads: SQO 2026-06-09 precedes SQM 2026-06-14.

If both flagged conversions and their pipeline are excluded pending correction, while retaining their SQMs, linkedin_ads has 8 − 2 = 6 valid SQOs: conversion = 6 ÷ 25 = 24%; cost/SQO = $24,000 ÷ 6 = $4,000; pipeline = $96,000 − 2 × $12,000 = $72,000; pipeline/$ = $72,000 ÷ $24,000 = 3x.

Reallocation recommendation and confidence

Pause paid_social pending attribution validation: $18,000 has no recorded SQMs. If tracking is complete, redirect that budget primarily to paid_search, which has the strongest observed pipeline/$ at 20x. Consider reducing linkedin_ads after correcting its dates. Preserve webinars as a measured test: its $1,800 cost/SQO is lowest, but rests on only 5 SQOs.

Confidence is moderate in the directional shift toward paid_search (40 SQMs, 18 SQOs), but low in an exact allocation or expected incremental return. webinars (12/5), linkedin_ads (25/8, including two invalid date sequences), organic_search (30/10), and referral (15/6) have limited samples. Attribution completeness, closed-won outcomes, and marginal performance at higher spend are missing; recent SQMs may also be immature. Pipeline/$ is not realized revenue ROI.
competitor-battlecard-updatePASS 1.00
Plain-English question: Update the Rivally battlecard with sourced pricing and where each side wins.
Prompt — what the model saw
Update the battlecard for the competitor. Sections: one-line positioning, pricing with source and date (newer source wins, note the conflict), where they win, where we win, objections and responses, recent changes, our 12-month win/loss record against them. Cite a snippet id for every factual claim. Rep opinion on a call is not a fact about the competitor. Anything from the old card you cannot re-source gets marked unverified.
cost $0.3808 · 51s · in 9 / out 1,856 tok
✓ competitor_named✓ pricing✓ win_loss✓ unverified_marked✓ citations

✓ Correct Answer (ground truth)

Plain-English answer: Rivally at $7 per user per month (the August source wins over the stale $5, conflict noted), a 13-7 win/loss record, and two unverified old claims flagged.
The answer in one line: competitor = Rivally
citation_ids_available25 items — e.g. S01, S02, S03, S04
competitorRivally
pricingvalue_per_user_per_month=7; source_snippet=S17; source_date=2026-08-12
rep_opinion_snippet_ids_not_factsS09, S21
unverified_old_claimsRivally lacks a Slack integration., Rivally was acquired by WorkHuman in 2025.
win_loss_12mowins=13; losses=7
graded checkscompetitor_named · pricing · win_loss · unverified_marked · citations

Model's actual answer

show the model's full answer
Battlecard: Rivally

One-line positioning
Points-based recognition with a reviewer-praised feed and multilingual support for distributed EU teams; EU data residency is generally available. [S02, S12, S15] The old card’s “for mid-market” positioning is unverified: one mid-market review does not establish target-market positioning. [S04]

Pricing — source, date, and conflict
- Current published price: Recognition Starter at $7/user/month, annual billing required — pricing page, 2026-08-12. [S17]
- Conflict: the old card’s $5/user/month, annual billing was supported by the 2026-01-20 pricing page; Starter still showed $5 on 2026-04-01. The newer pricing page wins: replace $5 with $7. Increase: $7 − $5 = $2/user/month; $2 ÷ $5 × 100 = 40%. [S03, S08, S17]
- Deal-specific evidence, not general pricing: call notes dated 2026-06-02 report $6.50/user/month for a 500-seat prospect on an annual term. [S13]
- A prospect reportedly received $7/user/month list with a 15% discount for a three-year term on 2026-08-14: $7 × (1 − 0.15) = $5.95/user/month. This is a reported offer, not a published discount policy. [S18]
- Rivally Pulse is a separately priced add-on, not bundled; no add-on amount is provided. [S23]

Where they win
- EU requirements: announced generally available EU data residency; an EU enterprise reviewer praised multilingual support and suitability for distributed EU teams. This supports the old card’s EU-strength claim with qualification, not a universal enterprise advantage. [S15, S12]
- Fast setup and Slack: one mid-market reviewer reported setup in under a week and Slack working out of the box. The old claim “Rivally lacks a Slack integration” is contradicted; remove it. [S04]
- Recognition experience: reviewers praised the points-based feed and its engagement. [S02, S16]
- Support: one reviewer praised response times under four hours; this is not evidence of a contractual SLA. [S22]

Where we win
- Analytics: an 800-seat prospect reportedly picked Bonusly over Rivally citing analytics depth. This is a specific reported win, not proof of universal superiority. [S25]
- Supporting competitive angles: reviewers described limited analytics and basic reporting dashboards. [S02, S07]
- Discovery areas, not established Bonusly advantages: reviewers reported missing SCIM, lagging admin tooling, and missing bulk recognition editing. No supplied evidence establishes Bonusly’s corresponding capabilities. [S10, S16, S24]

Objections and responses
- “Rivally is cheaper.”
  Response: “Its latest published Starter price is $7/user/month with annual billing, not the old $5. Compare equivalent scope and terms; no Bonusly pricing is supplied to establish which is cheaper.” [S17, S03, S08]
- “We need EU residency and multilingual support.”
  Response: “Rivally announced EU residency GA, and an EU enterprise reviewer praised multilingual support. Validate your specific requirements rather than denying these strengths.” [S15, S12]
- “Their analytics are sufficient.”
  Response: “Test your required reports and exports. Reviews describe basic reporting and CSV-only analytics exports; one prospect reportedly chose Bonusly for analytics depth.” These are attributed reports, not independently verified product limitations. [S07, S20, S25]
- “Their EMEA rewards coverage meets our needs.”
  Response: “Validate the catalog in your employees’ countries; one reviewer described EMEA coverage as thinner than the US catalog.” [S14]
- “Their administration will scale.”
  Response: “Ask for a demonstration of provisioning and bulk editing; reviewers reported absent SCIM and bulk recognition editing.” Current availability needs validation. [S10, S24]

Recent changes
- 2026-03-05: launched Rivally Pulse, a lightweight engagement-survey add-on; 2026-09-01: Pulse exited beta as a paid add-on, not bundled. [S06, S23]
- 2026-05-09: hired an ex-Workday VP EMEA to lead European expansion; 2026-07-01: opened a Dublin office and announced EU data residency GA. [S11, S15]
- 2026-08-12: published Starter pricing rose to $7/user/month from the earlier $5. [S17, S08]
- 2026-08-20: announced Microsoft Teams app v2 in public preview—not general availability. [S19]
- Old-card acquisition claim: “Rivally was acquired by WorkHuman in 2025” — UNVERIFIED; no supplied snippet supports it. S01 reports a $40M Series C led by Northgate Ventures, not an acquisition, and does not independently disprove one. [S01]
- Excluded as competitor facts: Bryce Harmon’s “clunky” UI assessment and Elena Sinclair’s aggressive-discounting impression are explicitly unconfirmed rep opinions. [S09, S21]

Our 12-month win/loss record
Period: September 2025–August 2026, the 12 months covered by deals_with_competitor.csv. This file has no snippet IDs; exact deal aliases below serve as row citations rather than inventing snippet IDs.

- Wins (13): Deal-A9FD43, Deal-7AA785, Deal-44C524, Deal-0D0CD6, Deal-D5B790, Deal-5C636E, Deal-67BE14, Deal-1B6969, Deal-F03E7B, Deal-072E31, Deal-F65C8F, Deal-E46EAB, Deal-1D2392.
- Losses (7): Deal-7767F5, Deal-5645A5, Deal-C6FFAA, Deal-D263E0, Deal-935746, Deal-9066A6, Deal-72A02F.
- Total = 13 + 7 = 20; win rate = 13 ÷ 20 × 100 = 65%; loss rate = 7 ÷ 20 × 100 = 35%. [deals_with_competitor.csv, rows cited above]
- The September 2026 prospect win is outside this window and has no deal alias linking it to the outcomes file; it is not added to the record. [S25]
nurture-sequence-diagnosisPASS 1.00
Plain-English question: What is broken in each email sequence — rates, overlaps, tracking errors?
Prompt — what the model saw
Review the sequences. Per sequence: sent, open rate, reply rate, meeting rate, weakest step. Identify tracking errors (opened above sent), audience overlap between sequences, and the failure mode for anything under 2 percent reply. One change per weak sequence and which to fix first. Under 350 words.
cost $0.6674 · 61s · in 21 / out 1,946 tok
✓ tracking_error✓ overlap_pair✓ weakest_step✓ fix_priority

✓ Correct Answer (ground truth)

Plain-English answer: Per-sequence open, reply, and meeting rates with the weakest step (all three weakest at step 3), an opened-above-sent tracking error, and audience overlap flagged.
The answer in one line: fix_priority_first = Cold Outbound - HR Leaders
all_pairwise_overlaps['Cold Outbound - HR Leaders', 'Cold Outbound - People Ops'], ['Expansion Nurture', 'New Logo Nurture']
audience_overlap_pairCold Outbound - HR Leaders, Cold Outbound - People Ops
failure_modeCold Outbound - HR Leaders=opens healthy (~37%) but replies near zero from step 1 -> body/CTA not landing, not a deliverability or subject-line problem
fix_priority_firstCold Outbound - HR Leaders
per_sequence{"New Logo Nurture": {"sent": 1386, "opened": 490, "replied": 90, "meetings": 27, "open_rate": 0.3535, "reply_rate": 0.0649, "meeting_rate": 0.0195, "weakest_step": 3}, "Expansion Nurture": {"sent": 8
tracking_errorsequence=Expansion Nurture; step=2; sent=300; opened=340; issue=opened exceeds sent
under_2pct_replyCold Outbound - HR Leaders
graded checkstracking_error · overlap_pair · weakest_step · fix_priority

Model's actual answer

show the model's full answer
Rates = summed events ÷ summed sends × 100, not unique-contact conversion. Weakest step = lowest reply rate.

Sequence results
• New Logo Nurture: sent = 500+458+428 = 1,386. Open: 490/1,386 = 35.35%; reply: 90/1,386 = 6.49%; meeting: 27/1,386 = 1.95%. Weakest: step 3, replies 18/428 = 4.21%.

• Expansion Nurture: sent = 300+300+275 = 875. Open: 565/875 = 64.57% (unreliable); reply: 59/875 = 6.74%; meeting: 12/875 = 1.37%. Weakest: step 3, replies 12/275 = 4.36%.

• Cold Outbound - HR Leaders: sent = 600+595+590 = 1,785. Open: 545/1,785 = 30.53%; reply: 8/1,785 = 0.45%; meeting: 0/1,785 = 0%. Weakest: step 3, replies 1/590 = 0.17%.

• Cold Outbound - People Ops: sent = 400+386+377 = 1,163. Open: 340/1,163 = 29.23%; reply: 29/1,163 = 2.49%; meeting: 6/1,163 = 0.52%. Weakest: step 3, replies 6/377 = 1.59%.

Tracking error
Expansion Nurture step 2: 340/300 = 113.33% opened; 340−300 = 40 above sent. Correct tracking before trusting opens; corrected count is missing.

Audience overlap
New Logo Nurture ↔ Expansion Nurture: 2 contacts—CT-000301, CT-000624.

Cold Outbound - HR Leaders ↔ Cold Outbound - People Ops: 21 contacts—CT-000849, CT-000884, CT-000890, CT-000908, CT-001033, CT-001097, CT-001101, CT-001103, CT-001105, CT-001130, CT-001153, CT-001159, CT-001217, CT-001227, CT-001236, CT-001255, CT-001258, CT-001277, CT-001285, CT-001311, CT-001345. No other overlaps in supplied rows.

Failure modes and changes
• Fix Cold Outbound - HR Leaders first: highest send volume, lowest reply rate, zero meetings. Every step is below 2%: 5/600=0.83%, 2/595=0.34%, 1/590=0.17%. Observed failure: opens without replies. Change: test a role-specific value proposition.
• Cold Outbound - People Ops: step 3 alone falls below 2%; late-step response deterioration. Change: replace step 3 with a distinct offer.

Copy, timing, and delivery data are missing; relevance, fatigue, and overlap effects remain hypotheses, not proven causes.
weekly-marketing-goals-rollupPASS 1.00
Plain-English question: Are we on pace on the quarter's marketing goals?
Prompt — what the model saw
Produce the weekly marketing goals update for the active quarter: SQMs, SQOs, DS2s, closed-lost MIA rate, same-quarter close count, active pipeline coverage against target. For each: QTD actual, target, delta, pace (on/behind/ahead given days elapsed). One paragraph on what moved this week.
cost $0.2494 · 24s · in 6 / out 734 tok
✓ metrics_present✓ deltas✓ coverage_ratio✓ pace✓ paragraph

✓ Correct Answer (ground truth)

Plain-English answer: SQMs ahead of pace, SQOs on pace, DS2s behind (35 short of expected at day 66 of 92), with deltas and same-quarter closes.
The answer in one line: days_in_quarter = 92; days_elapsed = 66; coverage_ratio_active_pipeline = 0.75; elapsed_fraction = 0.7174 (+1 more below)
days_in_quarter92
coverage_ratio_active_pipeline0.75
days_elapsed66
elapsed_fraction0.7174
metrics{"SQMs": {"qtd_actual": 230, "target": 300, "delta": -70, "expected_at_day_elapsed": 215.22, "pace": "ahead"}, "SQOs": {"qtd_actual": 84, "target": 120, "delta": -36, "expected_at_day_elapsed": 86.09,
pace_ruleexpected = target * days_elapsed/days_in_quarter; within +/-5% of expected = on, above = ahead, below = behind; lower_better metrics invert.
this_week_paragraph_required1
graded checksmetrics_present · deltas · coverage_ratio · pace · paragraph

Model's actual answer

show the model's full answer
Q3-2026 weekly marketing goals update

Quarter elapsed: 66 ÷ 92 = 71.74%. Pace uses a straight-line benchmark: target × 66/92. Delta = QTD actual − full-quarter target. MIA rate is compared directly with its ceiling, not prorated.

| Metric | QTD actual | Target | Delta | Pace |
|---|---:|---:|---:|---|
| SQMs | 230 | 300 | −70 | Ahead: 230 vs. 215.22 expected |
| SQOs | 84 | 120 | −36 | Behind: 84 vs. 86.09 expected |
| DS2s | 40 | 75 | −35 | Behind: 40 vs. 53.80 expected |
| Closed-lost MIA rate | 5 ÷ 25 = 20% | ≤10% | +10 percentage points (unfavorable) | Behind: twice the ceiling |
| Same-quarter close count | 10 | 20 | −10 | Behind: 10 vs. 14.35 expected |
| Active pipeline coverage | 3,000,000 ÷ 4,000,000 = 75% | 100% (4,000,000) | −25 percentage points (−1,000,000) | Ahead of linear benchmark: 75% vs. 71.74%; still below full target |

Active pipeline is a snapshot; its “ahead” classification assumes a linear build toward the quarter-end target, not demonstrated growth.

What moved this week: Weekly movement cannot be determined because no prior-week snapshot or weekly activity data was provided. The supplied QTD snapshot shows SQMs and active pipeline above the linear benchmark, while SQOs, DS2s, and same-quarter closes are behind; the MIA rate exceeds its ceiling. No specific weekly changes or drivers can be established.
communication 5 tests
ceo-slack-compression0.80
Plain-English question: Compress the forecast analysis into one Slack message the CEO can read in ten seconds.
Prompt — what the model saw
Compress the forecast analysis below into a single Slack message for the CEO. Rules: at most three sentences; under 80 words; lead with the decision or the so-what; keep every distinct number that matters and the one caveat that changes the conclusion; cut any sentence that does not carry a claim, a number, or a directive. No em dashes, no throat-clearing phrases.
cost $0.3652 · 22s · in 9 / out 430 tok
✓ max_3_sentences✓ under_80_words✓ weighted_present✗ caveat_present✓ no_em_dash

✓ Correct Answer (ground truth)

Plain-English answer: Three sentences, under 80 words, decision first, every distinct number kept plus the one caveat — all checked by deterministic text rules.
The answer in one line: weighted_number = 115976; must_keep = data-quality caveat — deleting the bad news fails
weighted_number115,976
limitmax 3 sentences, max 80 words, no em-dashes
must_keepdata-quality caveat — deleting the bad news fails
graded checksmax_3_sentences · under_80_words · weighted_present · caveat_present · no_em_dash

Model's actual answer

show the model's full answer
Treat Q3 2026’s 115,976.75 forecast as provisional: all 7 in-quarter COMMIT deals lack documented why-buys. Across 54/86 deals, 100% × 44,729 COMMIT + 35% × 203,565 BEST_CASE = 115,976.75; PIPELINE contributes 0. Exclude 32 deals totaling 227,575 closing after September 30, including Deal-D348E1 (13,770 COMMIT, October 15) and 9 BEST_CASE totaling 28,240; validate buying rationale before relying on the forecast.
stalled-deal-followup-emailPASS 1.00
Plain-English question: Write a follow-up email for this stalled deal — under 120 words, one prior touch, one ask.
Prompt — what the model saw
Write a follow-up email for this stalled deal. Under 120 words. Reference one specific prior touch from the engagement history, make one clear ask, and match the tone of the prior thread. No invented details.
cost $0.3224 · 16s · in 6 / out 203 tok
✓ word_cap✓ references_touch✓ single_ask

✓ Correct Answer (ground truth)

Plain-English answer: An email that references the July 28 milestone-awards demo, makes one clear ask, and invents nothing — Deal-0D2F7A, 31 days since last engagement.
The answer in one line: amount = 5100; word_cap = 120; days_since_last_engagement = 31; deal_alias = Deal-0D2F7A (+4 more below)
amount5,100
days_since_last_engagement31
word_cap120
allowed_facts14 items — e.g. Deal-0D2F7A, $5,100, DS3, Alex Franklin
allowed_newENG-00001, ENG-00002, ENG-00003, ENG-00004, ENG-00005
deal_aliasDeal-0D2F7A
must_referenceengagement_id=ENG-00002; date=2026-07-28; type=meeting; why=most substantive prior touch; concrete prospect reactions to milestone awards and points catalog
ownerAlex Franklin
single_cta1
stageDS3
tonewarm, concise, professional (matches prior thread)
graded checksword_cap · references_touch · single_ask

Model's actual answer

show the model's full answer
Subject: Deal-0D2F7A — pricing follow-up

Hi team,

I'm following up on my August 5 email with the July 28 demo recap and pricing for 150 seats.

Is this still something you'd like to move forward with?

Thanks,
Alex Franklin
funnel-digest-paragraphPASS 1.00
Plain-English question: Write the weekly GTM digest — one paragraph per team, every number traceable.
Prompt — what the model saw
Write the weekly GTM digest: one paragraph per team (Marketing, Sales, CS, PLG) in Bonusly brand voice. Every number must trace to the extracts. No customer contact names, no employee performance commentary, no compensation. Four paragraphs, under 400 words total.
cost $0.3527 · 20s · in 9 / out 346 tok
✓ four_paragraphs✓ word_cap✓ teams

✓ Correct Answer (ground truth)

Plain-English answer: Four paragraphs (Marketing, Sales, CS, PLG) using only the allowed numbers, no contact names, no performance commentary, under 400 words.
The answer in one line: word_cap = 400
word_cap400
allowed_numbers16 items — e.g. 2, 3, 9, 08
bannedcustomer contact names, employee performance commentary, compensation
required_paragraphsMarketing, Sales, CS, PLG
graded checksfour_paragraphs · word_cap · teams

Model's actual answer

show the model's full answer
Marketing: We brought in 46 SQMs against a weekly target of 52, leaving a gap of 6 (52 − 46 = 6). Webinars contributed 18 of those SQMs, a bright spot worth celebrating. Paid search cost per SQM held at $150.

Sales: We converted 14 SQOs and set 9 DS2 meetings this week. New pipeline totaled $310,000, and the same-quarter close count was 3. Those closes are worth a cheer, with new conversations on the calendar.

CS: We saved 2 renewals this week, a win worth recognizing. Team NPS moved to 61; the extract doesn’t provide the previous score, so the size and direction of the change aren’t available. Heading into next week, 3 open red-flag accounts remain on the radar.

PLG: We welcomed 412 new signups, with activation at 31%. And 38 companies reached the aha moment of 10 recognition gives—a milestone worth celebrating as more teams put appreciation into practice.
partner-digest-sectionPASS 1.00
Plain-English question: Write the partner digest so every partner in the roster appears exactly once.
Prompt — what the model saw
Write the partner digest for the period. Every partner in the roster appears once: active plays get status with numbers, quiet partners get one line. Partner-sourced pipeline count and amount must match the deal data. Nothing omitted, nothing added.
cost $0.2315 · 15s · in 6 / out 415 tok
✓ all_partners✓ pipeline_amount✓ quiet_one_line

✓ Correct Answer (ground truth)

Plain-English answer: Four active partners with status and numbers, four quiet partners with one line each, and partner-sourced pipeline that reconciles to the deal data.
The answer in one line: rule = every roster partner named exactly once; quiet partners get one line; no partners outside roster.
active_no_dealsWorkWell Group
partner_sourced_pipelinecount=5; amount=550000
per_partner_pipeline{"Apex Rewards Co": {"count": 2, "amount": 275000}, "HRCloud Partners": {"count": 1, "amount": 140000}, "CultureBridge": {"count": 2, "amount": 135000}}
quiet_partnersRecogniq, KudosWave, PeopleFirst Advisors, TotalPerk
rosterpartner_name=Apex Rewards Co; status=active | partner_name=HRCloud Partners; status=active | partner_name=CultureBridge; status=active | partner_name=WorkWell Group; status=active | partner_name=Recogniq; status=quiet | partner_name=KudosWave; status=quiet … (+2 more)
ruleevery roster partner named exactly once; quiet partners get one line; no partners outside roster.
graded checksall_partners · pipeline_amount · quiet_one_line

Model's actual answer

show the model's full answer
Partner digest | 2026-08-24 to 2026-09-04

- Apex Rewards Co — Co-webinar locked for 09-15. 2 sourced opportunities, both DS1: Deal-DDAAF2 (180,000) + Deal-2CF33E (95,000) = 275,000.
- HRCloud Partners — Security review completed; opportunity moved to DS2. 1 sourced opportunity: Deal-F1CDA5 (140,000).
- CultureBridge — Lunch-and-learn produced 2 sourced opportunities, both early stage: Deal-096E1D (60,000) + Deal-067213 (75,000) = 135,000.
- WorkWell Group — Joint playbook restart requested for Q4; planning call booked for 09-09. 0 sourced deals; 0 pipeline amount.
- Recogniq — Quiet; no activity reported or sourced deals listed.
- KudosWave — Quiet; no activity reported or sourced deals listed.
- PeopleFirst Advisors — Quiet; no activity reported or sourced deals listed.
- TotalPerk — Quiet; no activity reported or sourced deals listed.

Partner-sourced pipeline (UTM Source = Partner):
Count: 2 + 1 + 2 = 5 deals.
Amount: 180,000 + 95,000 + 140,000 + 60,000 + 75,000 = 550,000.
Currency not provided.
rfp-answer-from-policyPASS 1.00
Plain-English question: Answer the 10 RFP items from the policy excerpts only — no guessing.
Prompt — what the model saw
Answer the 10 questionnaire items from the policy excerpts only. Cite the excerpt id for every answer. Where the excerpts do not answer the question, mark it unanswerable and say what document would; do not guess. Tone: vendor response, direct.
cost $0.2668 · 15s · in 6 / out 372 tok
✓ answered_count✓ unanswerable_flagged✓ citations✓ no_guess

✓ Correct Answer (ground truth)

Plain-English answer: Seven answered with excerpt citations (Q1-Q7), three flagged unanswerable with the document that would answer them (Q8-Q10).
The answer in one line: answered_count = 7; rule = unanswerable items must be flagged and the document that would answer them named; do not guess.
answered_count7
answer_map{"Q1": ["P1"], "Q2": ["P1"], "Q3": ["P2"], "Q4": ["P3"], "Q5": ["P4"], "Q6": ["P5"], "Q7": ["P6"]}
ruleunanswerable items must be flagged and the document that would answer them named; do not guess.
unanswerableQ8, Q9, Q10
valid_citation_idsP1, P2, P3, P4, P5, P6, P7, P8
graded checksanswered_count · unanswerable_flagged · citations · no_guess

Model's actual answer

show the model's full answer
Q1. Yes. All customer data is encrypted at rest using AES-256. [P1]

Q2. Yes. Data in transit is encrypted using TLS 1.2 or higher. [P1]

Q3. Yes. Bonusly maintains a SOC 2 Type II report, renewed annually. Prospects can obtain the report under NDA. [P2]

Q4. Customer data is retained for 90 days after the contract ends, then deleted from production systems. [P3]

Q5. Our current subprocessor list is published at trust.bonusly.com. Customers receive 30 days’ notice before a new subprocessor is added; the notification method is not specified. [P4]

Q6. Yes. SAML 2.0 single sign-on is available on the Enterprise plan. [P5]

Q7. Production databases are backed up daily. Backups are encrypted and retained for 35 days. [P6]

Q8. Unanswerable from the provided excerpts: ISO 27001 certification status is not stated. P2 addresses SOC 2 only. Required document: a current ISO 27001 certificate or an official certification-status statement. [P2; no excerpt answers this question]

Q9. Unanswerable from the provided excerpts: no contractual uptime SLA percentage is stated. Required document: the contractual Service Level Agreement. [P1–P8: not addressed]

Q10. Unanswerable from the provided excerpts: willingness to sign a HIPAA Business Associate Agreement is not stated. Required document: an official HIPAA/BAA contracting policy or written confirmation from Legal. [P1–P8: not addressed]
ops-maintenance 5 tests
skill-file-reconciliationPASS 1.00
Plain-English question: Audit a skill set against its manifest and flag every conflict with a severity and an action.
Prompt — what the model saw
Reconcile the skill set below against its manifest. Find and report (with severity CRITICAL/WARNING/INFO and an action type TRIM_DESC, UPDATE_BODY, DELETE_SKILL, MERGE, or REVIEW): (1) any two skills whose ALWAYS-trigger phrases overlap/duplicate, (2) any circular delegation chain between skills (name it), (3) any delegation target that does not exist (dangling, name it), (4) any version conflict and which skill should survive, (5) how many manifest descriptions exceed 1,024 characters, (6) any hardcoded page ids, dates, or person names in skill bodies, (7) manifest drift in BOTH directions (files with no manifest row, and manifest rows with no file). Make one proposal per finding; rewrite nothing.
cost $1.4134 · 87s · in 9 / out 3,842 tok
✓ duplicate_trigger_found✓ circular_chain✓ dangling_target✓ oversized_desc_count✓ hardcoded_values✓ manifest_drift✓ severity_and_action

✓ Correct Answer (ground truth)

Plain-English answer: A severity-ranked list of TRIM_DESC / UPDATE_BODY / DELETE_SKILL / MERGE / REVIEW actions covering every rule violation found.
The answer in one line: files_audited = 14; rank = CRITICAL / WARNING / INFO with one fix action per issue
files_audited14
findingsduplicate trigger, circular delegation chain, dangling file reference, oversized description, hardcoded values, manifest drift
rankCRITICAL / WARNING / INFO with one fix action per issue
graded checksduplicate_trigger_found · circular_chain · dangling_target · oversized_desc_count · hardcoded_values · manifest_drift · severity_and_action

Model's actual answer

show the model's full answer
Scope: only the supplied manifest and skill files. “Dangling” means absent from this supplied set, not proven absent from an external installation.

1. Overlapping / duplicate triggers

| Severity | Action | Finding | Proposal |
|---|---|---|---|
| WARNING | MERGE | `comms-drafter` and `email-drafter` duplicate “write me an email,” “draft a follow-up,” “what should I say,” “bump email,” and “contract nudge,” plus email review/rewrite coverage. | Consolidate into `comms-drafter`, preserving `email-drafter`’s signature retrieval and context-sourcing behavior. This is a scope-based recommendation, not a verified version succession. |
| WARNING | TRIM_DESC | `pipeline-intelligence-report` and `weekly-pipeline-report` overlap on pipeline report/update requests and “what’s the pipeline look like” / “what does pipeline look like.” | Reserve scored/tiered active-deal reports for `pipeline-intelligence-report` and weekly funnel/target performance for `weekly-pipeline-report`. |
| WARNING | TRIM_DESC | `pipeline-intelligence-report` claims all pipeline questions and VP “forecast context”; `sales-forecast` claims “pipeline forecast,” quarter outlook, and deal-level confidence. | Narrow `pipeline-intelligence-report` to scoring requests, excluding quarter-revenue forecasts. |
| WARNING | TRIM_DESC | `pipeline-intelligence-report`’s blanket pipeline-question trigger includes `next-to-close` requests, although `next-to-close` explicitly says not to use it for shortlists. | Add an explicit shortlist exclusion to `pipeline-intelligence-report`. |
| WARNING | TRIM_DESC | `pipeline-intelligence-report`’s blanket trigger includes the stale/activity-gap requests assigned to `stale-pipeline-report`. | Exclude stale-contact and hygiene-only requests from its trigger. |
| WARNING | TRIM_DESC | `deal-strategy-coach` triggers on “which deals are likely to close”; `next-to-close` ALWAYS triggers on “which deals are most likely to close.” | Reserve immediate close shortlists for `next-to-close`; retain diagnosis/coaching in `deal-strategy-coach`. |
| WARNING | TRIM_DESC | `deal-strategy-coach` covers forecast risk, rep pipeline reviews, and stalled deals, overlapping `sales-forecast`, `pipeline-intelligence-report`, and `stale-pipeline-report`. | Limit coaching triggers to diagnostic/advisory intent rather than report-generation intent. |
| WARNING | TRIM_DESC | `deal-strategy-coach` triggers on “draft a manager email,” which also falls within both drafting skills’ broad triggers. | Make strategy requests the coaching entry point and route drafting-only requests to the consolidated drafter. |
| INFO | REVIEW | `model-selection` covers every task; `analysis-validator`, `signalforge-claim-compressor`, and `signalforge-feedback` overlap on analytical/report outputs. These are lifecycle stages, not necessarily duplicate task skills. | Review them as one ordered lifecycle rather than treating shared subject coverage as grounds for deletion. |

The drafting and coaching descriptions use “whenever” or “also trigger,” rather than the literal word ALWAYS; their unconditional trigger coverage still overlaps.

2. Circular delegation

CRITICAL · UPDATE_BODY — `deal-strategy-coach → email-drafter → deal-strategy-coach`.

Evidence:
- `deal-strategy-coach`, “Manager-to-prospect email frameworks”: use `email-drafter`.
- `email-drafter`, “Lane marker”: route strategic coaching to `deal-strategy-coach`.

Proposal: make the drafting handoff terminal for an already-diagnosed request, returning the draft to the caller without routing back. This is a conditional routing cycle, not proof that every invocation recurses.

No reverse delegation from `closed-lost-analysis` to `pipeline-intelligence-report` is specified. “Called from pipeline-intelligence-report” documents an inbound call, not a circular edge.

3. Dangling targets

Each row is a separate finding and proposal.

| Severity | Action | Missing target | Referencing skill(s) | Proposal |
|---|---|---|---|---|
| CRITICAL | REVIEW | `bonusly-data-questions` | `analysis-validator`, §12.4; also §G1-J | Verify and supply the specialist dependency before enabling its delegation. |
| CRITICAL | REVIEW | `bonusly-product-questions` | `analysis-validator`, §12.4 | Verify and supply the specialist dependency before enabling its delegation. |
| CRITICAL | REVIEW | `bonusly-business-reporting-questions` | `analysis-validator`, §12.4 | Verify and supply the specialist dependency before enabling its delegation. |
| CRITICAL | REVIEW | `bonusly-rewards-questions` | `analysis-validator`, §12.4 | Verify and supply the specialist dependency before enabling its delegation. |
| CRITICAL | REVIEW | `bonusly-ppp-questions` | `analysis-validator`, §12.4 | Verify and supply the specialist dependency before enabling its delegation. |
| CRITICAL | REVIEW | `bonusly-feature-flag-questions` | `analysis-validator`, §12.4 | Verify and supply the specialist dependency before enabling its delegation. |
| CRITICAL | REVIEW | `bonusly-deal-desk-questions` | `analysis-validator`, §12.4 | Verify and supply the specialist dependency before enabling its delegation. |
| CRITICAL | REVIEW | `bonusly-datadog-questions` | `analysis-validator`, §12.4 | Verify and supply the specialist dependency before enabling its delegation. |
| CRITICAL | REVIEW | `bonusly-brand` | `comms-drafter`, `email-drafter`, `sales-forecast`, `signalforge-claim-compressor` | Verify and supply the shared brand dependency. |
| CRITICAL | REVIEW | `prospect-research-multithreading` | `comms-drafter`, `email-drafter`, `deal-strategy-coach` | Verify and supply the contact-research handoff target. |
| CRITICAL | REVIEW | `signalforge-reports` | Mandatory skill-file reads in `pipeline-intelligence-report` and `weekly-pipeline-report` | Verify and supply the design skill and its referenced assets. |
| WARNING | REVIEW | `skill-orchestrator` | `analysis-validator` self-improvement references; `signalforge-feedback` registration checklist | Verify the missing integration dependency; these references do not establish a runtime delegation edge. |
| INFO | REVIEW | `caveman` | `signalforge-claim-compressor`, optional combined use | Verify availability before advertising the optional combination. |

Referenced supporting files such as `references/data-sources.md` were not supplied. Their existence cannot be determined from this input; they are not manifest skill rows.

4. Version conflicts and survivor

WARNING · UPDATE_BODY — `analysis-validator` declares v3.6 in its header/footer, but its validation-trail template identifies the validator as v3.2. Full Mode still lists G1-A–G1-H and G2-A–G2-E; the decision tree also omits G2-F from its Gate 2 range.

Proposal: retain `analysis-validator` v3.6 as authoritative and align active version labels/check ranges with it. Historical changelog entries should remain historical.

INFO · REVIEW — No competing versioned skill files are supplied. Neither `comms-drafter` nor `email-drafter` declares a version, so the data cannot establish that one is the newer release. `pipeline-intelligence-report` v6 using a v4 design system is not, by itself, a version conflict.

Proposal: do not delete a skill on version grounds without an explicit succession record.

5. Manifest descriptions exceeding 1,024 characters

INFO · REVIEW — 0 descriptions exceed the limit.

Arithmetic:
- Count = Σ[description_chars > 1,024] = 0 across 14 rows.
- Maximum = 1,006, shared by `pipeline-intelligence-report` and `signalforge-claim-compressor`.
- Headroom = 1,024 − 1,006 = 18 characters.

Proposal: retain all descriptions with respect to this length threshold; any trigger trimming above addresses routing, not oversize descriptions.

6. Hardcoded page IDs, dates, and person names in bodies

Each row groups literals within one skill into one finding and proposal. Historical/example literals are distinguished from operational dependencies.

| Severity | Action | Skill and evidence | Proposal |
|---|---|---|---|
| WARNING | UPDATE_BODY | `analysis-validator`: named Finance escalation to Manish / Amani; embedded roster including Alaina Loori, Shealagh Coughlin, Bryce Harmon, Hugo Lindqvist, Dana Mercer, Alex Franklin, Cole Ingram, Gavin Porter, Colleen Perry, Ellie Barton, Ashley Reyer, Megan Franz, Elena Sinclair, Youssef Elkhateeb, Amanda Czenkus, Ben Castelli, Amani Phipps, John Thomas, Yasmin Wahid; correction literals “Ashley Le” and “Tracy.” Dates include March 28, 2023; April 26, 2026; May 4, 2026; May 9, 2026; May 2026; Q1/Q2 2026 examples. | Resolve operational people/rosters dynamically and clearly separate dated historical evidence from current authority. |
| WARNING | UPDATE_BODY | `deal-strategy-coach`: playbook page `2257879045`, April 2026 playbook, “Pricing — 2026,” and routing to Perseus / Farid. | Move operational reference locations, pricing validity, and routing identities into maintained configuration. |
| WARNING | UPDATE_BODY | `partner-digest`: folder `2286616609`; pages `2286321666`, `2265382925`, `2236940297`, `2237825028`, `2239365136`, `2238283777`; Amani / Amani Phipps, Kelli, Jen Lee, Hani, Bryce, Sara; May 16, May 19, June 2, 2026, `2026-05-17`, and Q2/Q3 2026 references. | Parameterize destination/source pages and people, retaining dated issue references only as labeled examples/history. |
| WARNING | UPDATE_BODY | `sales-forecast`: parent page `2232582148`; Alaina in the operational Manager Forecast heading; Elena / Alaina in history; hardcoded “Open Q2 Deals” and “Q2 Narrative”; July 9, 2026 example and April 27, 2026 changelog. | Make operational quarter labels, leader identity, and destination configurable while retaining historical/example dates. |
| WARNING | UPDATE_BODY | `signalforge-feedback`: target page `2295136266`, parent `2234417154`, but activation checklist points to Build Log page `2247295002`; “Gavin Porter Rep Diagnostic” and “Q2 Pipeline Review” examples. | Consolidate operational logging references around one configured destination; keep person/quarter examples explicitly illustrative. |
| WARNING | UPDATE_BODY | `pipeline-intelligence-report`: fixed roster Bryce Harmon, Dana Mercer, Cole Ingram, Alex Franklin, Gavin Porter; May 2026 verification/version/deprecation references and March 2023 stale-table reference. | Replace the operational roster with live owner resolution and label dated system assertions as requiring re-verification. |
| WARNING | UPDATE_BODY | `weekly-pipeline-report`: Ben Lavin / Ben; fixed April 1–June 30, 2026 window, Q1 2026 context, and Q2/Q3 reporting buckets. Also fixed spreadsheet IDs `1CLZeOsElVDF_LF0ZG_t2nfwvhnZ6bpwqM_nX3WEYzcw` and `1ENuaEcCuLjdKhMvp8FK3Ys1ek5Aw9ZuOZhsHJJFoB_k`. | Parameterize recipient, reporting periods, and spreadsheet configuration. |
| INFO | REVIEW | `closed-lost-analysis`: May 2026 observations/field confirmations, MinIO vacation “May 4–12,” and illustrative “4/13” / “June 3.” | Review dated observations for continued relevance and prevent examples from becoming runtime dates. |
| INFO | REVIEW | `model-selection`: `last_checked: 2026-05-19`, dated changelog/deprecation assertions, knowledge-cutoff dates, and dated model identifier `claude-haiku-4-5-20251001`. | Preserve identifiers/history but verify the registry through its stated freshness procedure before using it. |
| INFO | REVIEW | `stale-pipeline-report`: example dates `5/15`, `5/19`, `5/7`; changelog `2026-06-10`. | Keep these explicitly illustrative/historical rather than inserting them into live outreach. |
| INFO | REVIEW | `signalforge-claim-compressor`: changelog `2026-05-09` and attribution handle `JuliusBrussee/caveman`. | Retain as provenance, not operational personalization. |

No hardcoded page IDs, fixed dates, or person names were identified in the bodies of `comms-drafter`, `email-drafter`, or `next-to-close`. The HubSpot `1973303` URL literal is a portal/account ID, not a page ID.

7. Manifest drift in both directions

| Severity | Action | Check | Arithmetic / result | Proposal |
|---|---|---|---|---|
| INFO | REVIEW | Supplied skill files without manifest rows | 14 files − 14 matched files = 0. | No manifest addition is indicated. |
| INFO | REVIEW | Manifest rows without supplied skill files | 14 rows − 14 matched rows = 0. | No manifest-row removal is indicated. |

All supplied file stems also match their manifest `declared_name` values. Missing referenced dependencies are dangling references, not evidence of files-without-rows drift.
runbook-from-slack-threadPASS 1.00
Plain-English question: Turn the resolved incident thread into a runbook.
Prompt — what the model saw
Turn the resolved incident thread into a runbook: ordered steps with the exact command or action taken, who ran it, how success was verified, and rollback for each step that changed state. Every step must trace to a message in the thread. Add nothing the thread does not contain; where the thread is unclear, mark the step as needing confirmation.
cost $0.3217 · 42s · in 9 / out 1,464 tok
✓ ordered_steps✓ rollback✓ needs_confirmation✓ sourced

✓ Correct Answer (ground truth)

Plain-English answer: Ordered steps from the thread, each with the exact command, who ran it, how success was verified, and rollback where state changed.
The answer in one line: rule = no steps beyond what the thread contains; every step traces to a message id.
ambiguous_steporder=4; msg=M05; mark=needs-confirmation
ruleno steps beyond what the thread contains; every step traces to a message id.
stepsorder=1; msg=M01; actor=Bryce Harmon; action=acknowledge PagerDuty alert, take IC; state_change=False; verified_by=M01 | order=2; msg=M02; actor=Farid Osman; action=check queue depth; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M02 (48,213 pending) | order=3; msg=M04; actor=Farid Osman; action=pause auto-enqueue via feature flag; command=bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'; state_change=True; rollback=bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'; verified_by=M07 (queue depth falling) | order=4; msg=M05; actor=Elena Sinclair; action=cleared the dead set; state_change=True; needs_confirmation=thread gives no command or method for clearing the dead set | order=5; msg=M06; actor=Bryce Harmon; action=scale reward-worker to 6 replicas; command=kubectl scale deployment/reward-worker --replicas=6; state_change=True; rollback=kubectl scale deployment/reward-worker --replicas=3; verified_by=M07/M08 | order=6; msg=M08; actor=Cole Ingram; action=verify queue drained and error rate baseline; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M08 (returns 0) … (+2 more)
graded checksordered_steps · rollback · needs_confirmation · sourced

Model's actual answer

show the model's full answer
1. Acknowledge the alert and take incident command — Bryce Harmon [M01]
   Action: Acknowledge PagerDuty alert for reward-worker queue depth > 10k; take IC.
   Verification: Acknowledgment and IC assignment stated in the message; separate verification not documented.
   Rollback: Not documented — needs confirmation.

2. Check queue depth — Farid Osman [M02]
   Command: `bundle exec rake sidekiq:queue_depth`
   Verification: Returned 48,213 pending reward jobs; normal is under 500.
   Rollback: Not applicable; read-only check.

3. Inspect the dead set — Farid Osman [M03]
   Action: Reported 112 dead jobs, all Redis::TimeoutError from around 13:58. Exact inspection command not documented — needs confirmation.
   Verification: Findings reported in the message.
   Rollback: Not applicable; inspection only.

4. Pause enqueue — Farid Osman [M04]
   Command: `bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'`
   Verification: No explicit confirmation that enqueue stopped — needs confirmation.
   Rollback: `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`

5. Clear the dead set — Elena Sinclair [M05]
   Action: “While I was in the console I cleared out the dead set.” Exact command not documented — needs confirmation.
   Verification: Completion reported; no verification result documented — needs confirmation.
   Rollback: Not documented — needs confirmation.

6. Scale workers from 3 to 6 — Bryce Harmon [M06]
   Command: `kubectl scale deployment/reward-worker --replicas=6`
   Verification: No replica-count verification documented — needs confirmation. Subsequent queue observations appear in [M07–M08], but do not directly verify the replica count.
   Rollback: `kubectl scale deployment/reward-worker --replicas=3`

7. Check queue progress — Farid Osman [M07]
   Action: Reported queue depth down to 9,400 and falling approximately 1,200/min. Exact measurement command not documented — needs confirmation.
   Verification: Queue depth and trend reported in the message.
   Rollback: Not applicable; observation only.

8. Verify queue drainage and error recovery — Cole Ingram [M08]
   Command: `bundle exec rake sidekiq:queue_depth`
   Action: Check error rate in Datadog.
   Verification: Command returned 0; Datadog error rate back to baseline. Exact Datadog query or view not documented — needs confirmation.
   Rollback: Not applicable; read-only checks.

9. Re-enable enqueue — Bryce Harmon [M09]
   Command: `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`
   Verification: 40 new jobs processed cleanly in the next 3 minutes.
   Rollback: Not documented for this step — needs confirmation.

10. Scale workers back to 3 and resolve — Bryce Harmon [M10]
    Command: `kubectl scale deployment/reward-worker --replicas=3`
    Verification: Queue stable at 0; incident declared resolved. Direct replica-count verification not documented — needs confirmation.
    Rollback: Not documented for this step — needs confirmation.
datadog-incident-timelinePASS 1.00
Plain-English question: Sequence the failure from the logs — what broke first, and what cascaded after it?
Prompt — what the model saw
From the log slice, sequence the failure: first error with timestamp, the cascade in order, the service and job involved, and the Datadog query that would confirm the first error. State what the logs do not show.
cost $0.2751 · 28s · in 6 / out 780 tok
✓ first_error_ts✓ service_named✓ cascade_ordered✓ query✓ gaps_stated

✓ Correct Answer (ground truth)

Plain-English answer: A reward-service Redis timeout at 14:01:12Z cascades to sidekiq at 14:01:40 and api-gateway at 14:03:05, with the Datadog query that confirms the first error.
The answer in one line: job_class_involved = RewardGiveJob
cascade_orderservice=reward-service; first_error_at=2026-09-03T14:01:12Z | service=sidekiq; first_error_at=2026-09-03T14:01:40Z | service=api-gateway; first_error_at=2026-09-03T14:03:05Z | service=web-app; first_error_at=2026-09-03T14:03:30Z
clean_servicespostgres
confirming_queryservice:reward-service status:error "@message:Redis::TimeoutError"
first_errortimestamp=2026-09-03T14:01:12Z; service=reward-service; message=Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s
job_class_involvedRewardGiveJob
logs_do_not_showroot cause of the Redis timeout, redis-server's own logs, infrastructure metrics (CPU/memory/network), any deploy or config change preceding the incident, who applied mitigation or when (no actor fields in logs)
graded checksfirst_error_ts · service_named · cascade_ordered · query · gaps_stated

Model's actual answer

show the model's full answer
First error
2026-09-03T14:01:12Z — reward-service:
“Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s”

This is the first observed error in the supplied slice, not proof of the underlying cause.

Cascade in timestamp order (2026-09-03, UTC)
1. 14:01:20 → 14:01:30 → 14:01:40 — reward-service reports retries exhausted for RewardGiveJob.
2. 14:01:40 — sidekiq reports RewardGiveJob failed with Redis::TimeoutError; retrying in 60s. The two 14:01:40 events cannot be ordered more precisely.
3. 14:01:46 — J-00005 is the earliest failed_at entry in sidekiq_jobs.csv, for RewardGiveJob. Further RewardGiveJob failures follow.
4. 14:02:28 — sidekiq reports another RewardGiveJob failure/retry.
5. 14:02:30 — sidekiq warns: “Queue reward depth above 10,000”.
6. 14:02:36 — RecognitionDigestJob also fails with Redis::TimeoutError, first recorded as J-00013.
7. 14:03:05 — api-gateway reports “502 upstream timeout calling reward-service /gives”.
8. 14:03:30 — web-app reports “Give form submission failed: upstream 502 from api-gateway”.
9. Through 14:06:52 — job failures/retries, gateway timeouts, and form-submission failures continue.
10. 14:22:10 — reward-service reports Redis connection restored and job processing resuming.
11. 14:24:45 — sidekiq reports reward queue depth below 500.

Service and job
The first failing service is reward-service; the explicitly named job is RewardGiveJob, with execution failures also logged by sidekiq. RecognitionDigestJob is additionally affected in sidekiq_jobs.csv.

Datadog confirmation query
Set the UTC time window to 2026-09-03T13:55:00Z–2026-09-03T14:02:00Z and sort oldest first:

```text
service:reward-service *:"Redis::TimeoutError" *:"redis-primary:6379" *:"timed out after 5s"
```

This query targets the first error’s service and message; it has not been executed here.

What the logs do not show
- Why Redis timed out, or whether failures began before this slice.
- Trace/job IDs linking the initial timeout to particular requests or failed jobs; chronology alone does not prove every causal link.
- A remediation action responsible for recovery.
- Exact queue peak, queue clearance, successful retries for individual jobs, or confirmed end-user recovery.
- Total affected users, lost/duplicated rewards, or business impact.
feature-flag-state-summaryPASS 1.00
Plain-English question: Summarize what each feature flag controls and who has it on.
Prompt — what the model saw
Summarize the state of the listed feature flags: what each controls per the code excerpt, which companies or segments have it on, targeting rules, and any flag with no code reference or no targets. Company counts must match the export. Do not describe a flag that is not in the export.
cost $0.2966 · 24s · in 9 / out 602 tok
✓ orphans✓ no_targets✓ counts_match✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: Recognition streaks on for 42 beta companies, redeem-flow redesign off with a 12-company target list, plus any flag with no code reference or no targets.
The answer in one line: rule = company counts must equal the export; do not describe flags absent from the export.
code_referencedrecognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, redeem_flow_redesign, analytics_dashboard_v3, ms_teams_app_v2
flags{"recognition_streaks_v2": {"state": "on", "targeting_rules": "segment:beta_companies", "company_count": 42}, "points_budget_guardrails": {"state": "on", "targeting_rules": "all_companies", "company_c
flags_with_no_targetspaused_offboard_cleanup
orphan_flags_no_code_referencelegacy_give_modal, survey_boosters_q3, paused_offboard_cleanup
rulecompany counts must equal the export; do not describe flags absent from the export.
graded checksorphans · no_targets · counts_match · no_invented

Model's actual answer

show the model's full answer
Source: flags_export.csv for state, targeting, and company counts; flag_code.md for behavior.

Company aliases and targeted-list members are not provided. Counts below are copied directly from the export, not calculated. Off flags retain configured targets; those counts do not mean the feature is on.

| Flag | State | Targeting rule | Export company_count | Controls per code excerpt |
|---|---|---|---:|---|
| recognition_streaks_v2 | on | segment:beta_companies | 42 | Records a give through `StreakTracker.record(give)` when enabled. |
| points_budget_guardrails | on | all_companies | 220 | Calls `BudgetService.new(company).enforce!(giver, points)` when enabled. |
| slack_dm_nudges | on | segment:region_na | 87 | Gates `SlackDm.send_nudge(user)`; the job returns without sending when disabled. |
| redeem_flow_redesign | off | targeted_list | 12 | Selects `RedeemV2Component` when enabled; otherwise renders `RedeemV1Component`. Configured targets are not currently on. |
| analytics_dashboard_v3 | on | segment:tier_three | 65 | Sets `@dashboard = AnalyticsV3.new(company)` when enabled. |
| ms_teams_app_v2 | off | targeted_list | 9 | Gates `TeamsAppV2.install(company)`. Configured targets are not currently on. |
| legacy_give_modal | off | segment:legacy_plan | 14 | No code reference in the provided excerpt; behavior unknown. Configured targets are not currently on. |
| survey_boosters_q3 | on | segment:legacy_plan | 7 | No code reference in the provided excerpt; behavior unknown. |
| paused_offboard_cleanup | off | Blank | 0 | No code reference in the provided excerpt; behavior unknown. No targets reported. |

Targeting gaps: segment definitions and evaluation logic are not provided. `legacy_give_modal` and `survey_boosters_q3` share `segment:legacy_plan` but have different exported counts—14 and 7, respectively; the provided data does not explain the difference.
nda-triagePASS 1.00
Plain-English question: Triage each NDA green, yellow, or red — and quote the clause that decides it.
Prompt — what the model saw
Triage each NDA as GREEN (standard approval), YELLOW (counsel review), or RED (full legal review). Cite the clause that drives each classification by quoting it. Catch embedded non-solicits, non-competes, missing carve-outs, and unusual term or governing law. One line of reasoning per NDA.
cost $0.2652 · 17s · in 6 / out 320 tok
✓ green✓ yellow✓ red✓ non_solicit_caught✓ clause_quoted✓ correct_labels

✓ Correct Answer (ground truth)

Plain-English answer: NDA-1 green (standard mutual terms), NDA-2 yellow for an embedded 18-month non-solicit, with the driving clause quoted for every classification.
The answer in one line: non_solicit_in = NDA-2
checksclassification_per_nda, clause_quotes_exist_in_document, non_solicit_caught, carveouts_checked
classifications{"NDA-1": {"classification": "GREEN", "driving_reason": "standard mutual NDA: 2-year term, standard carve-outs, Delaware law", "driving_clause_quote": "Carve-outs: information that (a) is or becomes p
non_solicit_inNDA-2
graded checksgreen · yellow · red · non_solicit_caught · clause_quoted · correct_labels

Model's actual answer

show the model's full answer
NDA-1 — GREEN (standard approval): Mutual confidentiality with all stated standard carve-outs and Delaware law; clause 2 sets a bounded duration: “Term: two (2) years from the Effective Date; obligations survive for three (3) years after termination.” No non-solicit or non-compete is stated.

NDA-2 — YELLOW (counsel review): Clause 4 embeds a non-solicit and hiring restriction beyond confidentiality: “During the term of this Agreement and for eighteen (18) months thereafter, neither party shall solicit for employment or hire any employee of the other party with whom it came into contact under this Agreement.”

NDA-3 — RED (full legal review): Clause 2 imposes a broad non-compete—“For a period of three (3) years following the Effective Date, Recipient shall not, directly or indirectly, engage in or provide services to any business that competes with Discloser's business.”—compounded by clause 4, “No confidentiality carve-outs are stated in this Agreement,” clause 1’s lengthy “Term: five (5) years,” and clause 3’s non-Delaware law and exclusive foreign forum: “the laws of the Republic of Ireland, and the parties submit to the exclusive jurisdiction of its courts.”