← back to all models ← RevenueBench home
stealth

stealth/space-bunny-alpha

Mean 0.942 · 28/38 perfect tests · $0.00 total · 38/40 tests completed — remaining tests discarded from study (endpoint timeouts after dedicated retries) · modalities in: text, image, video · out: text · each card: the prompt → the correct answer (gold) → this model's actual answer

What do the modalities mean?

Modalities in means what you can send this model: text, images, files, video, or audio. Modalities out means what it can send back. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what this model can handle in real work.

data-crm 5 tests
stage-hygiene-auditPASS 1.00
Plain-English question: How much of the open DS2+ pipeline is missing a substantive why-buys, and is anything missing DS1 timestamps or out of order?
Prompt — what the model saw
Audit stage hygiene on the open pipeline below. Rules: a deal at DS2 or later must have a substantive why-buys (15+ characters, not a placeholder such as tbd, n/a, see notes). A why_buys value of the form LONG(n) means the real text was present but truncated for display; treat it as present and substantive. Every deal should have a DS1 entry timestamp. A deal whose stage entry timestamps are out of order regressed (an earlier-stage timestamp is later than a later-stage timestamp). Report: percent of DS2+ deals non-compliant on why-buys overall and by owner (use the owners table for names), count of deals with no DS1 timestamp, count that regressed, and three process fixes. State definitions before numbers.
cost $0.0000 · 126s · in 110,540 / out 9,811 tok
✓ noncomp_pct✓ no_ds1_0✓ regressed_0✓ per_owner_breakdown✓ definitions_first✓ three_fixes

✓ Correct Answer (ground truth)

Plain-English answer: 81.5% of the 124 DS2+ deals are missing a substantive why-buys — Alex Franklin 48, Cole Ingram 19, Dana Mercer 18, Bryce Harmon 15, Farid Osman 1 — and every deal has a DS1 timestamp with none regressed.
The answer in one line: n_deals = 156; ds2plus = 124; noncomp_pct = 81.5; bryce_noncomp = 15; no_ds1 = 0; regressed = 0 (+1 more below)
n_deals156
bryce_noncomp15
ds2plus124
no_ds10
noncomp_pct81.5
regressed0
per_owner_noncompAlex Franklin=48 of 61 (78.7%); Cole Ingram=19 of 22 (86.4%); Dana Mercer=18 of 22 (81.8%); Bryce Harmon=15 of 18 (83.3%); Farid Osman=1 of 1 (100%)
graded checksnoncomp_pct · no_ds1_0 · regressed_0 · per_owner_breakdown · definitions_first · three_fixes

Model's actual answer

show the model's full answer
Definitions
- DS2+ = current stage DS2, DS3, DS4, or DS5.
- Why-buys non-compliant = blank, placeholder, or fewer than 15 substantive characters. Any LONG(n) value counts as present and substantive.
- Missing DS1 = blank t_ds1.
- Regressed = an earlier populated stage timestamp is later than a subsequent stage timestamp. Equal timestamps are not regressions.

Results
- Scope: 156 deals; 124 at DS2+ because 39 DS2 + 61 DS3 + 14 DS4 + 10 DS5 = 124.
- Overall why-buys non-compliance: 101 / 124 × 100 = 81.45%.
  - 23 deals were compliant, all represented as LONG(n); 101 were blank.

By owner
- Bryce Harmon: 15 / 18 × 100 = 83.3%
- Dana Mercer: 18 / 22 × 100 = 81.8%
- Alex Franklin: 48 / 61 × 100 = 78.7%
- Cole Ingram: 19 / 22 × 100 = 86.4%
- Farid Osman: 1 / 1 × 100 = 100.0%
- Elena Sinclair: N/A; 0 DS2+ deals

Check: 15 + 18 + 48 + 19 + 1 = 101 non-compliant; 18 + 22 + 61 + 22 + 1 = 124 DS2+ deals.

Other hygiene
- Deals with no DS1 timestamp: 0 / 156.
- Regressed deals: 0 / 156.

Three process fixes
1. Block DS2+ advancement until why_buys contains substantive text of at least 15 characters and is not a placeholder.
2. System-generate immutable stage-entry timestamps and validate that DS1 is populated and timestamps remain nondecreasing.
3. Create a weekly owner exception queue for blank why_buys fields, with completion reviewed during forecast inspection.
owner-resolution0.83
Plain-English question: Every open deal has a blank owner name — resolve each one to a real person using the owners table.
Prompt — what the model saw
The deal extract below has an owner name field that is blank on nearly every row and an owner id field that is populated. Using the owners table, resolve each open deal to an owner name. Report: number of deals resolved, deals whose owner id has no match in the owners table (list them), the owner ids that map to archived/deactivated owners (list their names), and the total pipeline amount per resolved owner. Do not guess a name for any unmatched id.
cost $0.0000 · 235s · in 119,312 / out 27,442 tok
✓ resolved_count✗ unmatched_none✓ deactivated_flagged✓ jake_amount✓ justin_amount✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: All 156 deals resolve to six named owners; jake_amount maps to Bryce Harmon ($1,054,144) and justin_amount to Alex Franklin ($624,310).
The answer in one line: jake_amount = 1,054,144; justin_amount = 624,310; resolved = 156
jake_amount1,054,144
justin_amount624,310
resolved156
deactivated
per_ownerBryce Harmon=1,054,144; Alex Franklin=624,310; Dana Mercer=341,195; Cole Ingram=288,161.43; Farid Osman=4,134; Elena Sinclair=2,100
graded checksresolved_count · unmatched_none · deactivated_flagged · jake_amount · justin_amount · no_invented

Model's actual answer

show the model's full answer
Resolution summary

- Deals resolved: 156 of 156.
  Arithmetic: 35 + 67 + 24 + 22 + 7 + 1 = 156.
- Deals whose owner_id has no owners.csv match: None.
- Archived/deactivated owner IDs used by open deals: None.
- Archived IDs present in owners.csv but not used by these deals:
  - 1520255671 — Gavin Porter
  - 77260721 — Hugo Lindqvist

Pipeline amount by resolved owner

Bryce Harmon (owner_id 119337721; 35 deals)

Deal-25F752=24,000 + Deal-E53952=19,656 + Deal-C26D20=13,500 + Deal-6787C2=7,000 + Deal-A5E80A=2,520
+ Deal-2D1F1B=240,000 + Deal-66D1FC=99,000 + Deal-C6FE92=72,000 + Deal-950043=70,000 + Deal-D73B89=63,600
+ Deal-B23205=45,000 + Deal-012CB1=1 + Deal-40522D=21,000 + Deal-C5658B=23,400 + Deal-523604=13,680
+ Deal-C9C286=5,502 + Deal-CA7DC0=8,160 + Deal-483B2D=1 + Deal-F0EBBB=11,400 + Deal-3795AD=1
+ Deal-332637=36,000 + Deal-1BEEBF=31,500 + Deal-E25A09=6,000 + Deal-FC22A3=10,800 + Deal-036E80=30,275
+ Deal-BB8880=17,400 + Deal-01E193=12,600 + Deal-C1FA6D=18,000 + Deal-7BBDFA=37,440 + Deal-A62B1D=18,828
+ Deal-333EBB=2,880 + Deal-93C8BF=36,000 + Deal-1CCE5C=20,880 + Deal-927338=10,920 + Deal-A414F6=25,200
= 1,054,144.00

Alex Franklin (owner_id 84342457; 67 deals)

Deal-5408B0=14,850 + Deal-D348E1=13,770 + Deal-547B2B=11,200 + Deal-403845=9,000 + Deal-A2B47C=6,360
+ Deal-C61CF7=5,400 + Deal-C6D97A=3,240 + Deal-F9A08A=2,484 + Deal-1FC049=1,920 + Deal-BA571A=1,080
+ Deal-3EED2C=7,200 + Deal-60C2C2=19,000 + Deal-FA053A=2,880 + Deal-7FA0C3=1,400 + Deal-E531A6=4,800
+ Deal-D0BC96=1,632 + Deal-5296C9=10,000 + Deal-885F45=9,300 + Deal-278DEC=2,700 + Deal-4A13AD=2,160
+ Deal-8AD4A5=1,800 + Deal-15D24F=3,600 + Deal-9D0060=3,840 + Deal-36C33F=15,000 + Deal-0D0211=1,968
+ Deal-5AD94B=4,000 + Deal-690476=3,600 + Deal-6C60D4=4,800 + Deal-EE195F=3,120 + Deal-F436DA=2,520
+ Deal-034D49=9,000 + Deal-6883F3=2,400 + Deal-EC3025=62,000 + Deal-317E6F=5,400 + Deal-0D2F7A=5,100
+ Deal-1E2498=16,700 + Deal-D1E6C2=4,400 + Deal-BE3D9D=1,620 + Deal-635B8E=2,600 + Deal-DCA846=7,200
+ Deal-D9A72E=18,000 + Deal-D9A12F=17,000 + Deal-C2FF3C=8,316 + Deal-CA5E44=8,100 + Deal-4F775F=18,000
+ Deal-898FC5=12,600 + Deal-CC08D1=24,000 + Deal-792D44=15,000 + Deal-293AF3=9,000 + Deal-D8ABF7=7,200
+ Deal-46988D=3,780 + Deal-E0B692=16,200 + Deal-712010=7,200 + Deal-13FEBD=4,680 + Deal-F67D31=1,800
+ Deal-E73427=18,000 + Deal-42F601=2,730 + Deal-ED725A=2,400 + Deal-55164C=3,060 + Deal-B936FE=18,000
+ Deal-4B0BEB=12,000 + Deal-D7E999=1,800 + Deal-819506=4,400 + Deal-530B50=31,200 + Deal-3BA5EA=7,200
+ Deal-5FDCE4=1,600 + Deal-92D97D=60,000
= 624,310.00

Dana Mercer (owner_id 83155923; 24 deals)

Deal-9AAE5F=11,250 + Deal-944310=10,500 + Deal-B7EBD1=9,000 + Deal-3974EB=9,000
+ Deal-2465CE=5,400 + Deal-62D607=4,800 + Deal-584EE5=4,600 + Deal-0660B4=1,920
+ Deal-57887A=15,000 + Deal-F336B6=4,200 + Deal-215CCA=18,900 + Deal-B42F46=27,000
+ Deal-E51FB7=43,875 + Deal-9DDE86=20,000 + Deal-44EA29=60,000 + Deal-F40F04=8,100
+ Deal-5EED42=16,250 + Deal-DAF1D9=3,150 + Deal-87DDD1=5,000 + Deal-8952F0=2,100
+ Deal-BA3DDC=23,400 + Deal-7E2131=5,400 + Deal-7599B8=7,350 + Deal-F9A3C1=25,000
= 341,195.00

Cole Ingram (owner_id 83155924; 22 deals)

Deal-7B3B0F=2,760 + Deal-FD9F4E=1,330 + Deal-AF932D=7,225.40 + Deal-C7F9BF=3,360
+ Deal-8BC9F5=5,616 + Deal-342E96=2,700 + Deal-FF809F=7,781.20 + Deal-A71728=6,947.50
+ Deal-B25F40=40,000 + Deal-CD47A6=12,168 + Deal-42326B=2,480.40 + Deal-FA32A0=11,116
+ Deal-627646=11,193 + Deal-E568D5=1,875 + Deal-1BA595=31,750 + Deal-813836=32,175
+ Deal-175395=4,779.88 + Deal-2F3A66=3,334.80 + Deal-D04904=58,529.25 + Deal-481E24=4,140
+ Deal-CFE1E8=18,000 + Deal-99A240=18,900
= 288,161.43

Farid Osman (owner_id 716654662; 7 deals)

Deal-499BF6=1,249 + Deal-03BA69=1 + Deal-117863=1 + Deal-F17780=1
+ Deal-8BA24E=2,880 + Deal-2B39B0=1 + Deal-8FDCD2=1
= 4,134.00

Elena Sinclair (owner_id 701163055; 1 deal)

Deal-57FF13=2,100
= 2,100.00

Reconciliation

1,054,144.00 + 624,310.00 + 341,195.00 + 288,161.43 + 4,134.00 + 2,100.00
= 2,314,044.43 total resolved pipeline amount.
gong-hubspot-join-integrity0.75
Plain-English question: How many call records point at deal ids that do not exist in the open pipeline?
Prompt — what the model saw
Check the integrity of the call-to-deal mapping below. The gong table lists calls in the last 90 days keyed to HubSpot deal ids; deals_open lists the current open default-pipeline deals. Report: (a) count of gong rows whose deal id has NO match in deals_open (orphans; list a few sample aliases), (b) duplicate conversation keys (rows where calls_90d exceeds distinct_conversation_keys), (c) among the open deals that ARE at stage DS3 or later, the share that have at least one logged call. Give sample aliases for the orphan finding, all taken from the data. Note: gong_owner_id is a Gong user id, not a HubSpot owner id.
cost $0.0000 · 82s · in 113,490 / out 9,395 tok
✓ orphan_count✗ dup_keys_0✓ ds3_coverage✓ sample_ids_present

✓ Correct Answer (ground truth)

Plain-English answer: 30 Gong call rows are orphans keyed to dead deal ids, none are duplicates, and DS3 call coverage is 29.4%.
The answer in one line: orphans = 30; ds3_cov_pct = 29.4; dup = 0
orphans30
ds3_cov_pct29.4
dup0
orphan_aliasesDeal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42, Deal-36EA09
graded checksorphan_count · dup_keys_0 · ds3_coverage · sample_ids_present

Model's actual answer

show the model's full answer
Join: gong hs_deal_id → deals_open deal_id. gong_owner_id was excluded from the match.

(a) Orphan Gong rows: 30 of 67 (44.8%)

Arithmetic: 67 total Gong rows − 37 matched rows = 30 orphans.  
30 ÷ 67 × 100 = 44.8%.

Sample orphan aliases:
- Deal-8FA85D
- Deal-8FC3F9
- Deal-3B7945
- Deal-42B265
- Deal-B038F0

(b) Rows where calls_90d exceeds distinct_conversation_keys: 0 of 67

Every row has calls_90d = distinct_conversation_keys.

(c) Open DS3+ deals with at least one logged call: 25 of 85 (29.4%)

Arithmetic:
- Total DS3+: 61 DS3 + 14 DS4 + 10 DS5 = 85
- With at least one call: 12 DS3 + 6 DS4 + 7 DS5 = 25
- Share: 25 ÷ 85 × 100 = 29.4%
snowflake-sql-trapped-catalog0.67
Plain-English question: Write one Snowflake query for first-month usage per customer using only the provided, trap-laden catalog.
Prompt — what the model saw
Using ONLY the catalog provided, write one Snowflake SQL query returning, per customer company, for its first calendar month as a customer: unique givers, recognition count, successful redemption count. Read the catalog notes carefully: some tables are stale or unpopulated and must be avoided, and one documented filter (excluding deleted givers) must NOT be applied to historical giving counts. Put a comment above each FROM or JOIN explaining the table choice. Return only fenced SQL.
cost $0.0000 · 197s · in 46,000 / out 19,737 tok
✓ avoids_stale_deals✓ avoids_unpopulated_live✗ uses_cohort_summary✗ succeeded_filter✓ no_giver_deletion_filter✓ returns_sql

✓ Correct Answer (ground truth)

Plain-English answer: A single correct SQL query that uses the right catalog tables and avoids the trap columns — graded on the exact tables and columns.
The answer in one line: use_table = PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
avoid_stalePRODUCTION.HUBSPOT.DEALS (stale)
avoid_unpopulatedV2_LIVE.OBJECTS_DEALS (unpopulated)
forbidden_filterdeleted-giver exclusion must NOT be applied
outputone fenced SELECT with per-FROM/JOIN table-choice comment
required_filtersucceeded redemptions only
use_tablePRODUCTION.PLG.COMPANY_COHORT_SUMMARY
graded checksavoids_stale_deals · avoids_unpopulated_live · uses_cohort_summary · succeeded_filter · no_giver_deletion_filter · returns_sql

Model's actual answer

show the model's full answer
```sql
-- NOT COMPUTABLE FROM THE PROVIDED CATALOG.
-- Required arithmetic:
--   unique givers          = distinct giver IDs during the company's first calendar month;
--   recognition_count      = recognition events during that month;
--   successful_redemptions = redemption events where STATE = 'succeeded' during that month.
-- PRODUCTION.PLG.COMPANY_COHORT_SUMMARY is limited to self-serve companies, and the catalog
-- does not document M1_* as calendar-month metrics or M1_USERS as unique givers.
-- PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 documents STATE = 'succeeded'
-- but exposes no documented company identifier, company name, or redemption event timestamp.
-- No recognition-event table with company, giver, and event timestamp fields is provided.
-- The catalog also provides no all-customer first-calendar-month field or company-level join path.
-- The deleted-giver exclusion is intentionally not applied because it must not be used for
-- historical giving counts.
-- Return no fabricated company-level results.
SELECT
    CAST(NULL AS VARCHAR) AS customer_company,
    CAST(NULL AS DATE) AS first_calendar_month,
    CAST(NULL AS BIGINT) AS unique_givers,
    CAST(NULL AS BIGINT) AS recognition_count,
    CAST(NULL AS BIGINT) AS successful_redemption_count,
    'NOT_COMPUTABLE_FROM_PROVIDED_CATALOG' AS data_status
WHERE FALSE;
```
crm-completeness-enrichmentPASS 1.00
Plain-English question: Audit the CRM for missing and invalid fields, and propose the fixes.
Prompt — what the model saw
Audit the CRM extract for completeness and propose fixes. Deals need owner, stage, amount, close date, why-buys; companies need industry, employee count, HQ country; contacts need email, title, persona. Report completeness percent per field. Identify duplicate company clusters (name variants, shared domain) with a survivor per cluster. List invalid emails and domain mismatches. Fill missing company fields only where the enrichment export has a matching row; where CRM and enrichment disagree list both and recommend a source. Never invent a value. End with the 10 fixes with the most pipeline amount at stake.
cost $0.0000 · 324s · in 38,766 / out 34,363 tok
✓ dup_clusters✓ invalid_emails✓ survivor_language✓ completeness_pct✓ no_invent

✓ Correct Answer (ground truth)

Plain-English answer: 9 companies missing employee count, 6 missing HQ country, 4 invalid emails, 1 domain mismatch, 2 duplicate clusters, and 25 enrichment matches.
The answer in one line: n_contacts = 52; n_companies = 34; enrichment_matches = 25; invalid_email_count = 4; n_dup_clusters = 2; domain_mismatch_count = 1 (+4 more below)
n_contacts52
domain_mismatch_count1
enrichment_matches25
invalid_email_count4
n_companies34
n_dup_clusters2
company_missingindustry=0; employee_count=9; hq_country=6
domain_mismatchesCT-0011
duplicate_clusters['C-0A092931', 'C-0A092932'], ['C-0A092933', 'C-0A092934']
invalid_emailsCT-0010, CT-0080, CT-0081, CT-0192
graded checksdup_clusters · invalid_emails · survivor_language · completeness_pct · no_invent

Model's actual answer

show the model's full answer
COMPLETENESS

Completeness = nonblank required fields ÷ total extract rows. Company percentages use all 34 company rows before deduplication.

Deals
| Required field | Populated | Completeness |
|---|---:|---:|
| owner | N/A | N/A |
| stage | N/A | N/A |
| amount | N/A | N/A |
| close date | N/A | N/A |
| why-buys | N/A | N/A |

No deals file or deal rows were supplied. This does not mean zero deals; deal completeness and pipeline amount cannot be calculated.

Companies
| Field | Populated | Missing | Completeness |
|---|---:|---:|---:|
| industry | 34/34 | 0 | 100.0% |
| employee_count | 25/34 | 9 | 73.5% |
| hq_country | 28/34 | 6 | 82.4% |

Companies with all three fields populated before enrichment:  
21/34 = 61.8%

Contacts
| Field | Populated | Missing | Completeness |
|---|---:|---:|---:|
| email, nonblank | 52/52 | 0 | 100.0% |
| title | 39/52 | 13 | 75.0% |
| persona | 37/52 | 15 | 71.2% |

Contacts with all three fields nonblank:  
30/52 = 57.7%

After excluding invalid or domain-mismatched emails, contacts with all three valid requirements:  
27/52 = 51.9%

MISSING RECORDS

Company employee_count — 9:
C-EC3025, C-96039F, C-44EA29, C-D04904, C-B23205, C-60C75F, C-7BBDFA, C-50D386, C-93C8BF

Company hq_country — 6:
C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB

Contact title — 13:
CT-0000, CT-0022, CT-0072, CT-0080, CT-0081, CT-0092, CT-0120, CT-0121, CT-0122, CT-0132, CT-0141, CT-0162, CT-0170

Contact persona — 15:
CT-0000, CT-0022, CT-0041, CT-0060, CT-0070, CT-0081, CT-0082, CT-0092, CT-0110, CT-0132, CT-0162, CT-0171, CT-0172, CT-0180, CT-0181

ENRICHMENT MATCHES AND ALLOWED FILLS

Matching company rows: 25/34 = 73.5%

Unmatched company aliases:
C-BA969B, C-332637, C-93C8BF, C-EE9FFB, C-C9BB20, C-0A092931, C-0A092932, C-0A092933, C-0A092934

Permitted missing-field fills:

| Company alias | Domain | Field | Enrichment value |
|---|---|---|---:|
| C-EC3025 | ec3025.com | employee_count | 400 |
| C-96039F | 96039f.com | employee_count | 400 |
| C-44EA29 | 44ea29.com | employee_count | 400 |
| C-D04904 | d04904.com | employee_count | 400 |
| C-B23205 | b23205.com | employee_count | 400 |
| C-60C75F | 60c75f.com | employee_count | 400 |
| C-7BBDFA | 7bbdfa.com | employee_count | 400 |
| C-50D386 | 50d386.com | employee_count | 400 |

No missing HQ country can be filled from enrichment: the five matching rows also have blank enrichment HQ values, and C-EE9FFB has no matching row.

Post-fill company completeness:

| Field | Before | After permitted fills |
|---|---:|---:|
| industry | 34/34 = 100.0% | 34/34 = 100.0% |
| employee_count | 25/34 = 73.5% | 33/34 = 97.1% |
| hq_country | 28/34 = 82.4% | 28/34 = 82.4% |
| all three fields | 21/34 = 61.8% | 27/34 = 79.4% |

The all-three increase is 21 + 6 = 27 because C-44EA29 and C-D04904 remain incomplete without HQ country.

DUPLICATE COMPANY CLUSTERS

No company-name field was supplied, so name-variant matching cannot be assessed. Shared-domain matching finds two clusters.

| Domain | Records | Survivor | Duplicate |
|---|---|---|---|
| acme-corp.com | C-0A092931; C-0A092932 | C-0A092931 (provisional) | C-0A092932 |
| globex.io | C-0A092933; C-0A092934 | C-0A092933 (provisional) | C-0A092934 |

acme-corp.com disagreement:
- C-0A092931: industry `Technology`, employee_count `500`, hq_country `US`
- C-0A092932: industry `tech`, employee_count `510`, hq_country `USA`
- No enrichment row exists; employee-count truth cannot be determined.

globex.io disagreement:
- C-0A092933: industry `SaaS`, employee_count `200`, hq_country `US`
- C-0A092934: industry `Technology`, employee_count `200`, hq_country `US`
- No enrichment row exists; industry truth cannot be determined.

Survivors are provisional record-identity choices, not confirmation that their values are correct. Preserve and review conflicting fields before merging.

INVALID EMAILS AND DOMAIN MISMATCHES

Structurally invalid emails — 4:
- CT-0010: `user0@`
- CT-0080: `user0@`
- CT-0081: `user1@`
- CT-0192: `user2@`

Syntactically valid email: 48/52 = 92.3%

Domain mismatch — 1:
- CT-0011: `user1@other-domain.com`; contact domain/company domain: `66d1fc.com`

Valid and domain-aligned emails: 47/52 = 90.4%

CRM–ENRICHMENT DISAGREEMENTS

Industry

| Company alias | Domain | CRM | ZoomInfo enrichment |
|---|---|---|---|
| C-66D1FC | 66d1fc.com | `tech` | `Computer Software` |
| C-EC3025 | ec3025.com | `Technology` | `Computer Software` |
| C-44EA29 | 44ea29.com | `tech` | `Computer Software` |
| C-92D97D | 92d97d.com | `Technology` | `Computer Software` |
| C-D04904 | d04904.com | `Technology` | `Computer Software` |
| C-77A95A | 77a95a.com | `Technology` | `Computer Software` |
| C-AA8DDA | aa8dda.com | `Technology` | `Computer Software` |
| C-B25F40 | b25f40.com | `Technology` | `Computer Software` |
| C-60C75F | 60c75f.com | `tech` | `Computer Software` |
| C-425E2A | 425e2a.com | `Tech` | `Computer Software` |

C-425E2A’s raw CRM value has trailing whitespace: `Tech `.

Recommendation: use the CRM industry taxonomy for current reporting, pending owner/source verification. Do not overwrite with `Computer Software` solely because it is more specific; the supplied data does not establish that the taxonomies are equivalent. Preserve the enrichment value for review.

HQ country

| Company alias | Domain | CRM | ZoomInfo enrichment |
|---|---|---|---|
| C-66D1FC | 66d1fc.com | `US` | `United States` |
| C-950043 | 950043.com | `US` | `United States` |
| C-EC3025 | ec3025.com | `USA` | `United States` |
| C-96039F | 96039f.com | `USA` | `United States` |
| C-77A95A | 77a95a.com | `US` | `United States` |
| C-B23205 | b23205.com | `US` | `United States` |
| C-E51FB7 | e51fb7.com | `USA` | `United States` |
| C-D0662E | d0662e.com | `US` | `United States` |
| C-425E2A | 425e2a.com | `USA` | `United States` |
| C-2D7423 | 2d7423.com | `USA` | `United States` |

Recommendation: retain CRM as the record source but normalize `US`, `USA`, and `United States` to one reporting label, `United States`. These are label differences rather than demonstrated geographic conflicts.

Employee-count disagreements:
- None among the 25 matching rows where both values are populated.

TOP 10 FIXES

Pipeline amount at stake cannot be ranked or quantified because no deal amounts, deal records, or deal-to-company associations were supplied. The following is the complete 10-fix remediation list, ordered by supplied-record exposure rather than invented pipeline dollars.

1. Supply the deal extract with owner, stage, amount, close date, why-buys, and deal-to-company associations; this is required for pipeline-dollar prioritization.
2. Backfill persona on 15 contacts.
3. Backfill title on 13 contacts.
4. Reconcile the 10 industry disagreements, using CRM pending verification.
5. Normalize the 10 HQ-country label disagreements to `United States` while preserving source values.
6. Fill employee_count for eight matching aliases from enrichment—C-EC3025, C-96039F, C-44EA29, C-D04904, C-B23205, C-60C75F, C-7BBDFA, C-50D386—and source C-93C8BF separately.
7. Source the six missing HQ countries; enrichment cannot fill C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, or C-EE9FFB.
8. Correct or replace the four invalid emails: CT-0010, CT-0080, CT-0081, and CT-0192.
9. Merge the duplicate acme-corp.com and globex.io clusters only after preserving and reviewing conflicting fields.
10. Verify CT-0011 because `user1@other-domain.com` does not match `66d1fc.com`.
deal-intelligence 3 tests
call-transcript-extractionPASS 1.00
Plain-English question: Extract the CRM write-back fields from each call transcript as JSON.
Prompt — what the model saw
For each transcript, extract CRM write-back fields as JSON: why-buys (prospect statements only), pain points, stakeholders from the speaker list, budget signal (prospect-stated or null), timeline signal, competitor mentioned (only if the prospect raised it), next step (only if explicitly agreed), objections, confidence. Never fill a field from what the rep said or hoped.
cost $0.0000 · 189s · in 33,799 / out 20,778 tok
✓ json_emitted✓ budget_signal✓ null_discipline✓ stakeholders✓ competitor_prospect_only

✓ Correct Answer (ground truth)

Plain-English answer: Per transcript: prospect-only why-buys, pain points, stakeholders, budget and timeline signals, and competitors — e.g. TX-001: '$40k earmarked', live before open enrollment.
The answer in one line: rep = Alex Franklin
allowed_newTX-001, TX-002, TX-003, TX-004, TX-005, TX-006, Achievers, Bucketlist, Nectar
per_transcript{"TX-001": {"why_buys": ["automating anniversary and birthday awards"], "pain_points": ["HR team of three cannot keep up with it manually", "track everything in a spreadsheet, and people slip through
repAlex Franklin
rep_sourced_nulls{"TX-003": ["budget_signal \u2014 only the rep mentioned pricing/budget"], "TX-004": ["next_step \u2014 rep proposed, prospect said 'no promises'"], "TX-006": ["budget_signal \u2014 rep decoy, prospec
graded checksjson_emitted · budget_signal · null_discipline · stakeholders · competitor_prospect_only

Model's actual answer

show the model's full answer
[
  {
    "transcript_id": "TX-001",
    "deal_alias": "Deal-CFE7F4",
    "why-buys": [
      {
        "speaker": "Prospect (VP People)",
        "statement": "The big win for us would be automating anniversary and birthday awards — our HR team of three cannot keep up with it manually."
      }
    ],
    "pain_points": [
      {
        "speaker": "Prospect (VP People)",
        "statement": "The big win for us would be automating anniversary and birthday awards — our HR team of three cannot keep up with it manually."
      },
      {
        "speaker": "Prospect (HR Admin)",
        "statement": "Right now we track everything in a spreadsheet, and people slip through the cracks."
      }
    ],
    "stakeholders": [
      "Prospect (VP People)",
      "Prospect (HR Admin)"
    ],
    "budget_signal": {
      "source": "prospect-stated",
      "amount": "$40k",
      "scope": "engagement tools",
      "period": "this fiscal year",
      "statement": "We have about $40k earmarked for engagement tools this fiscal year."
    },
    "timeline_signal": [
      {
        "speaker": "Prospect (VP People)",
        "statement": "Ideally we would have this live before open enrollment in November."
      }
    ],
    "competitor_mentioned": {
      "name": "Achievers",
      "speaker": "Prospect (VP People)",
      "statement": "We looked at Achievers last year, but it was too heavy for a team our size."
    },
    "next_step": {
      "agreed_by": "Prospect (VP People)",
      "statement": "Yes — let's do the security review on September 12."
    },
    "objections": [
      {
        "speaker": "Prospect (HR Admin)",
        "statement": "One concern: we need SSO and audit logs for IT to sign off."
      },
      {
        "speaker": "Prospect (VP People)",
        "statement": "We looked at Achievers last year, but it was too heavy for a team our size."
      }
    ],
    "confidence": {
      "level": "high",
      "basis": "Every populated field is directly supported by an explicit prospect statement or prospect agreement."
    }
  },
  {
    "transcript_id": "TX-002",
    "deal_alias": "Deal-70BB30",
    "why-buys": [
      {
        "speaker": "Prospect (Head of Total Rewards)",
        "statement": "We want to tie recognition to retention for our hourly workforce — regretted turnover there is over 30%."
      }
    ],
    "pain_points": [
      {
        "speaker": "Prospect (Head of Total Rewards)",
        "statement": "We want to tie recognition to retention for our hourly workforce — regretted turnover there is over 30%."
      }
    ],
    "stakeholders": [
      "Prospect (Head of Total Rewards)",
      "Prospect (CFO)"
    ],
    "budget_signal": {
      "source": "prospect-stated",
      "amount": "$25k",
      "scope": "pilot",
      "period": "this quarter",
      "statement": "Finance has approved a $25k pilot budget for this quarter."
    },
    "timeline_signal": [
      {
        "speaker": "Prospect (CFO)",
        "statement": "We want a decision by end of September."
      }
    ],
    "competitor_mentioned": null,
    "next_step": {
      "agreed_by": "Prospect (CFO)",
      "statement": "Yes — send the pilot agreement and we'll route it to legal this week."
    },
    "objections": [
      {
        "speaker": "Prospect (CFO)",
        "statement": "Integration with Workday has to be rock solid — that's my one condition."
      }
    ],
    "confidence": {
      "level": "high",
      "basis": "Every populated field is directly supported by an explicit prospect statement or prospect agreement."
    }
  },
  {
    "transcript_id": "TX-003",
    "deal_alias": "Deal-530B50",
    "why-buys": [
      {
        "speaker": "Prospect (People Ops Manager)",
        "statement": "We need to make recognition visible across our 12 retail locations."
      }
    ],
    "pain_points": [
      {
        "speaker": "Prospect (People Ops Manager)",
        "statement": "Store managers have zero budget autonomy for on-the-spot recognition today."
      }
    ],
    "stakeholders": [
      "Prospect (People Ops Manager)"
    ],
    "budget_signal": {
      "source": "prospect-stated",
      "amount": null,
      "statement": "Store managers have zero budget autonomy for on-the-spot recognition today."
    },
    "timeline_signal": [
      {
        "speaker": "Prospect (People Ops Manager)",
        "statement": "Honestly there's no rush on our side until Q1."
      }
    ],
    "competitor_mentioned": {
      "name": "Bucketlist",
      "speaker": "Prospect (People Ops Manager)",
      "statement": "My CEO used Bucketlist at her last company and liked it."
    },
    "next_step": {
      "agreed_by": "Prospect (People Ops Manager)",
      "statement": "Yes, let's schedule a call with our CEO — I'll send two times."
    },
    "objections": [
      {
        "speaker": "Prospect (People Ops Manager)",
        "statement": "Honestly there's no rush on our side until Q1."
      },
      {
        "speaker": "Prospect (People Ops Manager)",
        "statement": "The CEO has to be sold first — she decides anything people-related."
      }
    ],
    "confidence": {
      "level": "high",
      "basis": "Every populated field is directly supported by an explicit prospect statement or prospect agreement."
    }
  },
  {
    "transcript_id": "TX-004",
    "deal_alias": "Deal-180D02",
    "why-buys": [
      {
        "speaker": "Prospect (VP People)",
        "statement": "We want to consolidate three separate recognition tools into one."
      }
    ],
    "pain_points": [
      {
        "speaker": "Prospect (VP People)",
        "statement": "We're paying for three tools and none of them talk to our HRIS."
      }
    ],
    "stakeholders": [
      "Prospect (VP People)",
      "Prospect (IT Security Lead)"
    ],
    "budget_signal": {
      "source": "prospect-stated",
      "amount": null,
      "approval_threshold": "under $15k annually",
      "statement": "If it's under $15k annually, I can approve it without going to the board."
    },
    "timeline_signal": [
      {
        "speaker": "Prospect (IT Security Lead)",
        "statement": "Our procurement cycle runs six to eight weeks minimum."
      },
      {
        "speaker": "Prospect (IT Security Lead)",
        "statement": "The security review took three months for our last vendor — that's my hesitation."
      }
    ],
    "competitor_mentioned": null,
    "next_step": null,
    "objections": [
      {
        "speaker": "Prospect (IT Security Lead)",
        "statement": "The security review took three months for our last vendor — that's my hesitation."
      }
    ],
    "confidence": {
      "level": "high",
      "basis": "All populated fields are directly supported by prospect statements; no next step is recorded because the prospect did not explicitly agree to one."
    }
  },
  {
    "transcript_id": "TX-005",
    "deal_alias": "Deal-F8767A",
    "why-buys": [
      {
        "speaker": "Prospect (HR Director)",
        "statement": "Two things: automate service milestones, and give us analytics on recognition equity across departments."
      }
    ],
    "pain_points": [
      {
        "speaker": "Prospect (People Ops Coordinator)",
        "statement": "Our night-shift teams feel invisible — their engagement scores run 20 points lower."
      }
    ],
    "stakeholders": [
      "Prospect (HR Director)",
      "Prospect (People Ops Coordinator)"
    ],
    "budget_signal": {
      "source": "prospect-stated",
      "amount": "$12k",
      "scope": "engagement line",
      "statement": "We have $12k approved under our engagement line."
    },
    "timeline_signal": [
      {
        "speaker": "Prospect (HR Director)",
        "statement": "We need this running before our January all-hands."
      },
      {
        "speaker": "Prospect (HR Director)",
        "statement": "Yes — come present to our exec team on October 2."
      }
    ],
    "competitor_mentioned": {
      "name": "Nectar",
      "speaker": "Prospect (HR Director)",
      "statement": "We're mid-pilot with Nectar right now, so you'd need to beat that experience."
    },
    "next_step": {
      "agreed_by": "Prospect (HR Director)",
      "statement": "Yes — come present to our exec team on October 2."
    },
    "objections": [
      {
        "speaker": "Prospect (HR Director)",
        "statement": "We're mid-pilot with Nectar right now, so you'd need to beat that experience."
      },
      {
        "speaker": "Prospect (HR Director)",
        "statement": "Our exec team is skeptical after a failed rollout two years ago."
      }
    ],
    "confidence": {
      "level": "high",
      "basis": "Every populated field is directly supported by an explicit prospect statement or prospect agreement."
    }
  },
  {
    "transcript_id": "TX-006",
    "deal_alias": "Deal-EE195F",
    "why-buys": [
      {
        "speaker": "Prospect (HR Manager)",
        "statement": "I'd love to cut the admin time on service awards."
      }
    ],
    "pain_points": [
      {
        "speaker": "Prospect (HR Manager)",
        "statement": "I personally spend five hours a month ordering and shipping plaques."
      }
    ],
    "stakeholders": [
      "Prospect (HR Manager)"
    ],
    "budget_signal": {
      "source": "prospect-stated",
      "amount": null,
      "statement": "Budget isn't the issue — time is."
    },
    "timeline_signal": [
      {
        "speaker": "Prospect (HR Manager)",
        "statement": "Our fiscal year resets in February, so a Q1 start is realistic."
      }
    ],
    "competitor_mentioned": null,
    "next_step": {
      "agreed_by": "Prospect (HR Manager)",
      "statement": "Yes — send the one-page overview and I'll forward it to our COO this week."
    },
    "objections": [
      {
        "speaker": "Prospect (HR Manager)",
        "statement": "Nobody else — we're comparing this against just doing it internally."
      },
      {
        "speaker": "Prospect (HR Manager)",
        "statement": "Fair warning, our COO usually prefers building things in-house."
      }
    ],
    "confidence": {
      "level": "high",
      "basis": "Every populated field is directly supported by an explicit prospect statement or prospect agreement."
    }
  }
]
next-to-close0.75
Plain-English question: Which 3 deals are closest to signature right now?
Prompt — what the model saw
Which 3 deals are closest to signature right now? Use the CRM data and the Slack thread; Slack often has the real status before the CRM does. For each: deal, amount, why it is close, what is left. Under 250 words.
cost $0.0000 · 47s · in 32,993 / out 3,898 tok
✓ exactly_three✗ top3_named✓ slack_first✓ blocked_excluded

✓ Correct Answer (ground truth)

Plain-English answer: Deal-547B2B ($11,200), Deal-B7EBD1 ($9,000), and Deal-403845 ($9,000) — all DS5 COMMIT deals closing September 10-11, with what is left on each.
blocked_excludeddeal_alias=Deal-2465CE; reason=champion left + procurement freeze (Slack #sales-sparring 2026-09-04)
exactly_three1
ranking_ruleSlack signature-imminent deal ranks first; remaining ranked by stage (DS5>DS4), forecast (COMMIT>BEST_CASE>PIPELINE), close_date asc, amount desc
top3rank=1; deal_alias=Deal-547B2B; amount=11200; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11 | rank=2; deal_alias=Deal-B7EBD1; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-10 | rank=3; deal_alias=Deal-403845; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11
graded checksexactly_three · top3_named · slack_first · blocked_excluded

Model's actual answer

show the model's full answer
Based only on the latest supplied evidence. Slack ends September 4, with no later signature confirmation, so “right now” cannot be independently verified.

1. Deal-547B2B | $11,200
Why close: DS5/COMMIT. Redlines returned clean, the signing page was sent, and the VP People said on September 4 that they were “signing tomorrow.” Arithmetic: September 4 + 1 day = September 5 planned signature.
What’s left: Buyer execution. No confirmation that it occurred.

2. Deal-403845 | $9,000
Why close: DS5/COMMIT. The order form is with finance, and Slack says it is “moving fine.”
What’s left: Finance’s action and execution. No completion status is provided.

3. Deal-A2B47C | $6,360
Why close: DS5/COMMIT. Slack calls it warm and at normal legal-review pace. Its September 11 close date is earlier than Deal-D348E1’s October 15 date among deals carrying the same Slack status.
What’s left: Legal review through completion and signature. No specific blocker or execution date is provided.

Amount check: $11,200 + $9,000 + $6,360 = $26,560.
product-gap-detectionPASS 1.00
Plain-English question: Which prospect complaints are true product gaps versus rollout or enablement issues?
Prompt — what the model saw
Surface true product gaps raised by prospects in active deals. For each candidate: quote the prospect line, classify as real gap, plan gate, or rollout/enablement issue using the product docs excerpt, and name the deal and amount. Only real gaps go in the summary. Rep language does not count as prospect voice.
cost $0.0000 · 53s · in 36,220 / out 4,728 tok
✓ real_gaps✓ plan_gate_classified✓ enablement_classified✓ gap_summary

✓ Correct Answer (ground truth)

Plain-English answer: Two real gaps (SCIM provisioning, ADP Workforce Now integration) separated from plan gates and enablement issues, each with the prospect quote and deal.
allowed_newTG-001, TG-002, TG-003, TG-004, TG-005, SCIM user provisioning, ADP Workforce Now integration, Custom report builder
enablement_issuestranscript_id=TG-004; deal_alias=Deal-84DBA6; amount=16000; feature=Slack integration recognition adoption; quote=We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.; why_enablement=Slack integration exists on all plans; managers were never trained
excluded_rep_voicetranscript_id=TG-005; line=Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap.; reason=rep voice, not prospect-raised
plan_gatestranscript_id=TG-003; deal_alias=Deal-CFE7F4; amount=32000; feature=Custom report builder; tier_required=Enterprise; quote=I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?
real_gapstranscript_id=TG-001; deal_alias=Deal-EC3025; amount=62000; feature=SCIM user provisioning; quote=We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.; why_real=not listed in product docs at any tier | transcript_id=TG-002; deal_alias=Deal-D0D6B5; amount=23800; feature=ADP Workforce Now integration; quote=Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.; why_real=supported HRIS list is Workday/BambooHR/Gusto/Rippling only
graded checksreal_gaps · plan_gate_classified · enablement_classified · gap_summary

Model's actual answer

show the model's full answer
Candidates

1. Real gap — Deal-EC3025
Prospect: “We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.”
Product basis: SCIM user provisioning is not listed as a supported capability.
Deal amount: Not provided.

2. Real gap — Deal-D0D6B5
Prospect: “Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.”
Product basis: ADP Workforce Now is not listed as a supported integration.
Deal amount: Not provided.

3. Plan gate — Deal-CFE7F4
Prospect: “I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?”
Product basis: The custom report builder is available only on Enterprise, not Core or Pro.
Deal amount: Not provided.

4. Rollout/enablement issue — Deal-84DBA6
Prospect: “We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.”
Product basis: Slack integration is available on all plans. The stated problem is adoption caused by missing training, not missing product capability.
Deal amount: Not provided.

Excluded: Deal-36C33F has no prospect-voiced gap. The native-mobile-app statement is from Alex Franklin, and the prospect says the web version is sufficient.

Summary — real product gaps only

- 2 real gaps: SCIM user provisioning for Deal-EC3025 and ADP Workforce Now for Deal-D0D6B5.
- Classification arithmetic: 2 real gaps + 1 plan gate + 1 rollout/enablement issue = 4 candidates.
- Deal amounts were not provided, so no amount-at-risk total can be calculated.
rep-performance 5 tests
stale-pipeline-by-rep0.67
Plain-English question: Which open deals have had no email, call, or meeting in the last 7 days?
Prompt — what the model saw
The snapshot date is 2026-09-05. Using the open deals and the engagements-by-deal table, list every open deal with no logged email, call, or meeting in the last 7 days. Compute recency from the engagements table's last_email/last_call/last_meeting fields, NOT from any date field on the deal (the deal's last_contacted_field is unreliable). For each stale deal give: deal alias, owner name, stage, amount, days since last contact. Group by owner, order by amount descending within each owner. Then state, per owner, how many stale deals they carry and the total stale amount. (Do not write outreach text.)
cost $0.0000 · 360s · in 96,560 / out 43,480 tok
✓ stale_recall_ge_80✓ grouped_by_owner✓ jake_top_amount✓ kevin_stale_amt✗ jenn_stale_amt✗ stale_count

✓ Correct Answer (ground truth)

Plain-English answer: 65 open deals are stale; the largest is Deal-2D1F1B at $240,000, with Bryce Harmon carrying the most stale dollars at $626,243.
The answer in one line: jake_amt = 626,243; jenn_amt = 261,645; kevin_amt = 252,905.03; top_amount = 240,000; stale_count = 65; top_alias = Deal-2D1F1B
jake_amt626,243
jenn_amt261,645
kevin_amt252,905.03
stale_count65
top_amount240,000
top_aliasDeal-2D1F1B
graded checksstale_recall_ge_80 · grouped_by_owner · jake_top_amount · kevin_stale_amt · jenn_stale_amt · stale_count

Model's actual answer

show the model's full answer
Method: A deal is stale when the latest last_email, last_call, or last_meeting dated on or before 2026-09-05 is before 2026-08-30. Days since contact = 2026-09-05 − latest qualifying date.

Future-dated last_meeting values after 2026-09-05 were excluded because they had not occurred by the snapshot date. The deals file’s last_contacted_field was not used.

Bryce Harmon

Deal alias       | Stage | Amount    | Days since contact
-----------------|-------|-----------|------------------
Deal-2D1F1B      | DS1   | 240000.00 | 81
Deal-66D1FC      | DS1   | 99000.00  | 16
Deal-950043      | DS1   | 70000.00  | 19
Deal-B23205      | DS1   | 45000.00  | 16
Deal-7BBDFA      | DS3   | 37440.00  | 46
Deal-332637      | DS2   | 36000.00  | 9
Deal-1BEEBF      | DS1   | 31500.00  | 19
Deal-A414F6      | DS1   | 25200.00  | 19
Deal-C5658B      | DS1   | 23400.00  | 16
Deal-40522D      | DS3   | 21000.00  | 19
Deal-C1FA6D      | DS1   | 18000.00  | 16
Deal-01E193      | DS1   | 12600.00  | 8
Deal-F0EBBB      | DS3   | 11400.00  | 24
Deal-927338      | DS1   | 10920.00  | 18
Deal-E25A09      | DS1   | 6000.00   | 9
Deal-C9C286      | DS2   | 5502.00   | 9
Deal-012CB1      | DS1   | 1.00      | 23
Deal-3795AD      | DS2   | 1.00      | 8

Stale count: 18

Stale amount:
240000 + 99000 + 70000 + 45000 + 37440 + 36000 + 31500 + 25200 + 23400 + 21000 + 18000 + 12600 + 11400 + 10920 + 6000 + 5502 + 1 + 1
= 692964.00

Dana Mercer

Deal alias       | Stage | Amount    | Days since contact
-----------------|-------|-----------|------------------
Deal-44EA29      | DS2   | 60000.00  | 10
Deal-E51FB7      | DS2   | 43875.00  | 12
Deal-B42F46      | DS1   | 27000.00  | 19
Deal-BA3DDC      | DS3   | 23400.00  | 15
Deal-9DDE86      | DS2   | 20000.00  | 15
Deal-215CCA      | DS3   | 18900.00  | 17
Deal-5EED42      | DS3   | 16250.00  | 11
Deal-57887A      | DS2   | 15000.00  | 8
Deal-944310      | DS4   | 10500.00  | 33
Deal-B7EBD1      | DS5   | 9000.00   | 16
Deal-3974EB      | DS4   | 9000.00   | 8
Deal-F40F04      | DS2   | 8100.00   | 15
Deal-7599B8      | DS3   | 7350.00   | 18
Deal-87DDD1      | DS1   | 5000.00   | 19
Deal-F336B6      | DS3   | 4200.00   | 15
Deal-0660B4      | DS4   | 1920.00   | 16

Stale count: 16

Stale amount:
60000 + 43875 + 27000 + 23400 + 20000 + 18900 + 16250 + 15000 + 10500 + 9000 + 9000 + 8100 + 7350 + 5000 + 4200 + 1920
= 279495.00

Cole Ingram

Deal alias       | Stage | Amount    | Days since contact
-----------------|-------|-----------|------------------
Deal-D04904      | DS2   | 58529.25  | 11
Deal-B25F40      | DS3   | 40000.00  | 8
Deal-813836      | DS2   | 32175.00  | 11
Deal-1BA595      | DS2   | 31750.00  | 11
Deal-CFE1E8      | DS3   | 18000.00  | 11
Deal-CD47A6      | DS2   | 12168.00  | 11
Deal-627646      | DS3   | 11193.00  | 11
Deal-FF809F      | DS2   | 7781.20   | 11
Deal-AF932D      | DS2   | 7225.40   | 11
Deal-A71728      | DS2   | 6947.50   | 11
Deal-8BC9F5      | DS2   | 5616.00   | 10
Deal-175395      | DS3   | 4779.88   | 11
Deal-481E24      | DS3   | 4140.00   | 10
Deal-C7F9BF      | DS2   | 3360.00   | 11
Deal-2F3A66      | DS3   | 3334.80   | 11
Deal-342E96      | DS2   | 2700.00   | 24
Deal-E568D5      | DS3   | 1875.00   | 11
Deal-FD9F4E      | DS5   | 1330.00   | 10

Stale count: 18

Stale amount:
58529.25 + 40000 + 32175 + 31750 + 18000 + 12168 + 11193 + 7781.20 + 7225.40 + 6947.50 + 5616 + 4779.88 + 4140 + 3360 + 3334.80 + 2700 + 1875 + 1330
= 252905.03

Alex Franklin

Deal alias       | Stage | Amount   | Days since contact
-----------------|-------|----------|------------------
Deal-CC08D1      | DS1   | 24000.00 | 16
Deal-E73427      | DS3   | 18000.00 | 10
Deal-885F45      | DS2   | 9300.00  | 12
Deal-C2FF3C      | DS1   | 8316.00  | 10
Deal-0D2F7A      | DS3   | 5100.00  | 12
Deal-6C60D4      | DS3   | 4800.00  | 12
Deal-13FEBD      | DS2   | 4680.00  | 12
Deal-819506      | DS1   | 4400.00  | 8
Deal-9D0060      | DS3   | 3840.00  | 12
Deal-690476      | DS2   | 3600.00  | 18
Deal-C6D97A      | DS4   | 3240.00  | 8
Deal-EE195F      | DS3   | 3120.00  | 8
Deal-278DEC      | DS3   | 2700.00  | 8
Deal-635B8E      | DS3   | 2600.00  | 18
Deal-6883F3      | DS1   | 2400.00  | 16
Deal-4A13AD      | DS3   | 2160.00  | 26
Deal-F67D31      | DS2   | 1800.00  | 8
Deal-5FDCE4      | DS3   | 1600.00  | 12
Deal-BA571A      | DS4   | 1080.00  | 18

Stale count: 19

Stale amount:
24000 + 18000 + 9300 + 8316 + 5100 + 4800 + 4680 + 4400 + 3840 + 3600 + 3240 + 3120 + 2700 + 2600 + 2400 + 2160 + 1800 + 1600 + 1080
= 106736.00

Farid Osman

Deal alias       | Stage | Amount   | Days since contact
-----------------|-------|----------|------------------
Deal-8BA24E      | DS1   | 2880.00  | 8
Deal-8FDCD2      | DS1   | 1.00     | 15

Stale count: 2

Stale amount:
2880 + 1
= 2881.00

Overall

Stale count:
18 + 16 + 18 + 19 + 2 = 73

Total stale amount:
692964.00 + 279495.00 + 252905.03 + 106736.00 + 2881.00
= 1334981.03

Missing engagement data

These open deals have no matching row in engagements_by_deal_90d.csv, so stale status and days since contact cannot be determined. They are excluded from the 73-deal stale count and 1334981.03 total.

Alex Franklin:
Deal-3EED2C | DS2 | 7200.00 | days unavailable

Elena Sinclair:
Deal-57FF13 | DS1 | 2100.00 | days unavailable

Engagement coverage: 154 of 156 open deals matched; 2 missing.
activity-mix-vs-outcomePASS 1.00
Plain-English question: Per rep, what did the last 30 days of emails, calls, and meetings look like next to the deals that advanced?
Prompt — what the model saw
The snapshot date is 2026-09-05. For each sales rep (owner), compute the last-30-day totals of emails, calls, and meetings from the engagements table, and the number of deals that entered DS2 in the last 30 days (t_ds2 within 30 days of the snapshot). Report per rep: the activity mix as percentages (emails/calls/meetings share of that rep's total activities), total activities per DS2 entry, and rank the reps by that efficiency ratio (lowest activities per DS2 entry = most efficient). Name the most efficient rep and the highest-volume rep, and note if they differ. Do not attribute cause.
cost $0.0000 · 223s · in 99,119 / out 15,601 tok
✓ alex_ds2_30d✓ justin_volume_high✓ efficient_named✓ mix_present✓ no_causal_claim

✓ Correct Answer (ground truth)

Plain-English answer: Alex Franklin is the standout signal with 18 deals entering DS2 in 30 days (next closest: Bryce Harmon with 4), alongside per-rep activity totals.
The answer in one line: alex_ds2_30d = 18; bryce_ds2_30d = 4; note = legacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
alex_ds2_30d18
bryce_ds2_30d4
notelegacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
graded checksalex_ds2_30d · justin_volume_high · efficient_named · mix_present · no_causal_claim

Model's actual answer

show the model's full answer
Snapshot: 2026-09-05  
DS2 window: 2026-08-06 through 2026-09-05 inclusive.

Activity total = `emails_30d + calls_30d + meetings_30d`. Inbound emails were not added separately.

| Rep | Activity arithmetic | Activity mix: Email / Call / Meeting | DS2 entries | Activities per DS2 entry | Rank |
|---|---:|---:|---:|---:|---:|
| Alex Franklin | 307 + 36 + 41 = 384 | 79.95% / 9.38% / 10.68% | 18 | 384 ÷ 18 = 21.33* | 1 |
| Bryce Harmon | 162 + 0 + 43 = 205 | 79.02% / 0.00% / 20.98% | 4 | 205 ÷ 4 = 51.25 | 2 |
| Cole Ingram | 96 + 14 + 1 = 111 | 86.49% / 12.61% / 0.90% | 2 | 111 ÷ 2 = 55.50 | 3 |
| Farid Osman | 38 + 0 + 34 = 72 | 52.78% / 0.00% / 47.22% | 1 | 72 ÷ 1 = 72.00 | 4 |
| Dana Mercer | 84 + 18 + 11 = 113 | 74.34% / 15.93% / 9.73% | 1 | 113 ÷ 1 = 113.00 | 5 |
| Elena Sinclair | Not computable: no engagement row for her only deal | Not computable | 0 | Not computable | Not ranked |

DS2-entry aliases counted:

- Alex Franklin, 18: `Deal-403845`, `Deal-1FC049`, `Deal-3EED2C`, `Deal-7FA0C3`, `Deal-E531A6`, `Deal-5296C9`, `Deal-36C33F`, `Deal-EE195F`, `Deal-F436DA`, `Deal-317E6F`, `Deal-D1E6C2`, `Deal-D9A72E`, `Deal-CA5E44`, `Deal-4F775F`, `Deal-898FC5`, `Deal-46988D`, `Deal-E73427`, `Deal-92D97D`
- Bryce Harmon, 4: `Deal-25F752`, `Deal-D73B89`, `Deal-CA7DC0`, `Deal-1CCE5C`
- Cole Ingram, 2: `Deal-42326B`, `Deal-1BA595`
- Farid Osman, 1: `Deal-499BF6`
- Dana Mercer, 1: `Deal-57887A`
- Elena Sinclair, 0: none

Most efficient rep: Alex Franklin, at 21.33 observed activities per DS2 entry.  
Highest-volume rep: Alex Franklin, with 384 observed activities.  
They are the same rep.

Data limitation:
- `Deal-3EED2C` has no row in `engagements_by_deal_90d.csv`. Alex Franklin’s 384 total and 21.33 ratio are therefore lower bounds based on the available engagement rows; the exact activity total for that deal is missing, so the ranking is provisional.
- `Deal-57FF13` also has no engagement row. Elena Sinclair’s activity total and efficiency ratio cannot be computed.
- Archived owners Gavin Porter and Hugo Lindqvist have no rows in `deals_open.csv`; no activity or DS2 result is computable for them from the provided files.
ae-qtd-scorecardPASS 1.00
Plain-English question: Build Alex Franklin's QTD scorecard: bookings versus quota, pipeline, and conversion.
Prompt — what the model saw
Build a QTD scorecard for the named AE as of 2026-09-05: bookings vs quota with attainment percent, new vs expansion split, active pipeline by stage with amount, rolling 90-day DS2-to-won rate, win and loss counts with top loss reason, activity volume by type in the last 30 days. Exclude any closed-won deal dated before the quarter. Close with three coaching observations grounded in the numbers.
cost $0.0000 · 455s · in 99,556 / out 30,379 tok
✓ bookings✓ quota✓ attainment✓ new_expansion_split✓ top_loss_reason✓ coaching_obs

✓ Correct Answer (ground truth)

Plain-English answer: 75% attainment ($150,000 on a $200,000 quota), $113,500 new versus $36,500 expansion, pipeline by stage, and the rolling DS2-to-won rate.
The answer in one line: quota = 200000; bookings = 150000; new_bookings = 113500; expansion_bookings = 36500; open_deals_count = 125; attainment_pct = 75; owner = Alex Franklin (+12 more below)
quota200,000
attainment_pct75
bookings150,000
expansion_bookings36,500
expansion_count3
losses_qtd27
new_bookings113,500
new_count5
open_deals_count125
wins_qtd8
activity_30d_totalsemails_30d=807; calls_30d=112; meetings_30d=128; notes_30d=50
allowed_new18 items — e.g. Deal-A1C3E5, Deal-B7D2F4, Deal-C9E1A6, Deal-D4B8C2
ds2_to_wonwindow_start=2026-06-07; won_in_window=8; closed_in_window=35; rate=0.2286
ownerAlex Franklin
pipeline_by_stage{"DS1": {"count": 20, "amount": 284621}, "DS2": {"count": 28, "amount": 353760}, "DS3": {"count": 67, "amount": 552705}, "DS4": {"count": 5, "amount": 23574}, "DS5": {"count": 5, "amount": 45730}}
prior_quarter_excludeddeal_alias=Deal-B3E6F1; close_date=2026-06-20; amount=24000
quarter2026-Q3
snapshot_date2026-09-05
top_loss_reasonreason=Lost- Timing (1 year or more); count=13
graded checksbookings · quota · attainment · new_expansion_split · top_loss_reason · coaching_obs

Model's actual answer

show the model's full answer
Alex Franklin — QTD scorecard as of 2026-09-05

Scope: 2026-Q3, close dates from 2026-07-01 through 2026-09-05. Amounts use the file’s unspecified currency; no currency or company aliases were provided.

BOOKINGS VS QUOTA

Q3 won deals:

New:
- Deal-A1C3E5 — 40,000
- Deal-B7D2F4 — 35,000
- Deal-C9E1A6 — 21,000
- Deal-D4B8C2 — 11,000
- Deal-E6F3A9 — 6,500
- New total: 40,000 + 35,000 + 21,000 + 11,000 + 6,500 = 113,500

Expansion:
- Deal-F2C7D8 — 20,000
- Deal-A8B4D6 — 12,000
- Deal-C5D9E2 — 4,500
- Expansion total: 20,000 + 12,000 + 4,500 = 36,500

Total bookings:
- 113,500 + 36,500 = 150,000
- Quota: 200,000
- Attainment: 150,000 ÷ 200,000 = 75.0%
- Remaining gap: 200,000 − 150,000 = 50,000

New vs expansion:
- New: 113,500 ÷ 150,000 = 75.7% of booking amount; 5 ÷ 8 = 62.5% of wins
- Expansion: 36,500 ÷ 150,000 = 24.3% of booking amount; 3 ÷ 8 = 37.5% of wins

Deal-B3E6F1 — 24,000, won 2026-06-20 — was excluded because it predates Q3.

ACTIVE PIPELINE

Definition: all 125 rows with status=open in the supplied file.

Stage        Deals   Amount
DS1             20    284,621
DS2             28    353,760
DS3             67    552,705
DS4              5     23,574
DS5              5     45,730
Total          125  1,260,390

Arithmetic:
284,621 + 353,760 + 552,705 + 23,574 + 45,730 = 1,260,390

The total includes one overdue open deal: Deal-7A2454 — DS3, 1,275, close date 2026-09-04. The other 124 open deals have close dates after the scorecard date.

ROLLING 90-DAY DS2-TO-WON RATE

Using a 14-day outcome-maturity buffer:
- Cutoff: 2026-09-05 − 14 days = 2026-08-22
- Window start: 2026-08-22 − 90 days = 2026-05-24
- Cohort: 8 won + 27 lost + 57 active = 92 deals
- DS2-to-won rate: 8 ÷ 92 = 8.7%

The cohort’s deal_type is blank for 84 of 92 records, so the rate uses all supplied deal types; a new-business-only rate cannot be calculated reliably.

WIN/LOSS — QTD

- Wins: 8
- Losses: 27
- Total decided deals: 8 + 27 = 35
- Count-based win rate: 8 ÷ 35 = 22.9%
- Top loss reason: Lost- Timing (1 year or more) — 13 of 27 losses = 48.1%

ACTIVITY — LAST 30 DAYS

Activity totals across all 161 deal-level engagement rows:
- Emails: 807
- Calls: 112
- Meetings: 128
- Notes: 50
- Total: 807 + 112 + 128 + 50 = 1,097

The file supplies 30-day aggregates but no engagement dates. No task, LinkedIn, or other activity-type fields were provided.

COACHING OBSERVATIONS

1. Closing pace: The remaining 50,000 gap across 25 calendar days after September 5 requires 50,000 ÷ 25 = 2,000 per day before any further slippage.

2. Funnel maturity: Active pipeline is 1,260,390 ÷ 200,000 = 6.3× quota, but 638,381 ÷ 1,260,390 = 50.6% is in DS1/DS2 and only 69,304 ÷ 1,260,390 = 5.5% is in DS4/DS5. Prioritize qualification and stage advancement over adding unconverted volume.

3. Activity quality: Emails account for 807 ÷ 1,097 = 73.6% of activity, while timing represents 13 ÷ 27 = 48.1% of losses. Use more live qualification and explicitly documented next steps to address timing losses rather than relying on email volume alone.
multithreading-gapPASS 1.00
Plain-English question: Which open deals are single-threaded or under-threaded?
Prompt — what the model saw
Find every open deal that is single-threaded (fewer than 2 active contacts) or under-threaded (fewer than 3, or all contacts in one persona). Active means engaged in the last 60 days and not marked former. Personas: economic buyer, champion, HR admin, IT security, finance. For each flagged deal: amount, stage, active contact count, personas present, personas missing, the most valuable persona to add given the stage, and any on-file unengaged contact at that company who fits; otherwise say none on file.
cost $0.0000 · 369s · in 51,711 / out 45,261 tok
✓ flagged_recall✓ persona_language✓ single_thread_term✓ on_file_fit

✓ Correct Answer (ground truth)

Plain-English answer: Deals with fewer than 2 active contacts (single-threaded) or fewer than 3, or all in one persona (under-threaded), each flagged with the missing persona.
The answer in one line: active_window_start = 2026-07-07
active_window_start2026-07-07
allowed_new59 items — e.g. C-077A0E, C-0D15DF, C-1BDFE0, C-2AE3AA
best_persona_ruleDS2/DS3: economic buyer > IT security > finance > HR admin > champion; DS4/DS5: economic buyer > finance > IT security > HR admin > champion; first persona not present among active contacts
flag_ruleflag if active_contacts < 2 (single-threaded) or active_contacts < 3 or all active contacts share one persona (under-threaded); active = engaged within 60 days and not former
flaggeddeal_alias=Deal-36C33F; amount=15000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-50D386; amount=36000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5408B0; amount=14850; stage=DS4; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5BFE3B; amount=18000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-885F45; amount=9300; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=IT security | deal_alias=Deal-92D97D; amount=60000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer … (+5 more)
healthy_not_flaggedDeal-4B0BEB, Deal-84DBA6, Deal-D348E1
persona_universeeconomic buyer, champion, HR admin, IT security, finance
graded checksflagged_recall · persona_language · single_thread_term · on_file_fit

Model's actual answer

show the model's full answer
SCOPE AND ARITHMETIC

Analysis date: 2026-09-24
60-day cutoff: 2026-09-24 − 60 days = 2026-07-26

Active contact test:
`last_engaged_date >= 2026-07-26 AND is_former = false`

Flag test:
`active contacts < 2 OR active contacts < 3 OR all active contacts are in one persona`

The supplied files contain 14 deal IDs. 11 meet the flag test:
`5 single-threaded + 4 with two active contacts + 2 with three active contacts in one persona = 11 flagged`.

Open/closed status, amount, and stage are not provided in either file. Open status therefore cannot be verified. Amount and stage are shown as “not provided.” A stage-based “most valuable persona” cannot be determined from the supplied data. Unengaged-fit recommendations below are based only on missing active personas.

FLAGGED DEALS

1. Deal-EC3025 | Company C-FDD0C7
   Amount: not provided
   Stage: not provided
   Active contacts: 1/2
   Flag: single-threaded; CT-F2C1AE excluded as former
   Personas present: champion
   Personas missing: economic buyer; HR admin; IT security; finance
   Most valuable persona to add: not determinable; stage missing
   Unengaged fit on file: CT-6827DB | Chief People Officer | economic buyer

2. Deal-92D97D | Company C-E23238
   Amount: not provided
   Stage: not provided
   Active contacts: 1/2
   Flag: single-threaded; CT-A902AE last engaged 2026-06-01, 115 days ago
   Personas present: HR admin
   Personas missing: economic buyer; champion; IT security; finance
   Most valuable persona to add: not determinable; stage missing
   Unengaged fit on file: none on file

3. Deal-50D386 | Company C-EB10E4
   Amount: not provided
   Stage: not provided
   Active contacts: 2/2
   Flag: under-threaded; fewer than 3 active contacts
   Personas present: champion; HR admin
   Personas missing: economic buyer; IT security; finance
   Most valuable persona to add: not determinable; stage missing
   Unengaged fit on file: CT-A1C4B3 | Chief People Officer | economic buyer

4. Deal-D0D6B5 | Company C-32918E
   Amount: not provided
   Stage: not provided
   Active contacts: 3/3
   Flag: all active contacts are in one persona; champion
   Personas present: champion
   Personas missing: economic buyer; HR admin; IT security; finance
   Most valuable persona to add: not determinable; stage missing
   Unengaged fit on file: CT-1FA4DB | Chief People Officer | economic buyer

5. Deal-5BFE3B | Company C-535D36
   Amount: not provided
   Stage: not provided
   Active contacts: 2/2
   Flag: under-threaded; fewer than 3 active contacts
   Personas present: champion
   Personas missing: economic buyer; HR admin; IT security; finance
   Most valuable persona to add: not determinable; stage missing
   Unengaged fit on file: none on file

6. Deal-36C33F | Company C-077A0E
   Amount: not provided
   Stage: not provided
   Active contacts: 1/3
   Flag: single-threaded; CT-405B45 and CT-86B22F excluded as former
   Personas present: IT security
   Personas missing: economic buyer; champion; HR admin; finance
   Most valuable persona to add: not determinable; stage missing
   Unengaged fit on file: CT-1DB73E | Chief People Officer | economic buyer

7. Deal-885F45 | Company C-5E8EFB
   Amount: not provided
   Stage: not provided
   Active contacts: 2/2
   Flag: under-threaded; fewer than 3 active contacts
   Personas present: economic buyer; champion
   Personas missing: HR admin; IT security; finance
   Most valuable persona to add: not determinable; stage missing
   Unengaged fit on file: CT-B3F25D | IT Security Lead | IT security

8. Deal-FCBE5B | Company C-737030
   Amount: not provided
   Stage: not provided
   Active contacts: 1/1
   Flag: single-threaded
   Personas present: champion
   Personas missing: economic buyer; HR admin; IT security; finance
   Most valuable persona to add: not determinable; stage missing
   Unengaged fit on file: none on file

9. Deal-5408B0 | Company C-2AE3AA
   Amount: not provided
   Stage: not provided
   Active contacts: 2/2
   Flag: under-threaded; fewer than 3 active contacts
   Personas present: champion; HR admin
   Personas missing: economic buyer; IT security; finance
   Most valuable persona to add: not determinable; stage missing
   Unengaged fit on file: CT-07FA76 | Chief People Officer | economic buyer

10. Deal-C6D97A | Company C-5A8FC2
    Amount: not provided
    Stage: not provided
    Active contacts: 3/3
    Flag: all active contacts are in one persona; champion
    Personas present: champion
    Personas missing: economic buyer; HR admin; IT security; finance
    Most valuable persona to add: not determinable; stage missing
    Unengaged fit on file: none on file

11. Deal-F9A08A | Company C-0D15DF
    Amount: not provided
    Stage: not provided
    Active contacts: 1/2
    Flag: single-threaded; CT-913581 last engaged 2026-06-20, 96 days ago
    Personas present: champion
    Personas missing: economic buyer; HR admin; IT security; finance
    Most valuable persona to add: not determinable; stage missing
    Unengaged fit on file: CT-697541 | Chief People Officer | economic buyer
talk-track-from-transcriptsPASS 1.00
Plain-English question: From this rep's last 10 calls: how do they open, handle objections, and how often do they lock a next step?
Prompt — what the model saw
From this rep's last 10 calls: what they lead with in the first five minutes, how they handle the three most common objections, how often a concrete next step is agreed (rate), and every competitor a prospect raised. Quote at most one line per finding and only lines that appear in the transcripts. End with two coaching notes.
cost $0.0000 · 129s · in 37,559 / out 12,538 tok
✓ objection_count✓ next_step_rate✓ competitors✓ coaching✓ quote_used

✓ Correct Answer (ground truth)

Plain-English answer: 8 of 10 calls open with the same customer-win story; objection handling, next-step agreement rate, and every competitor raised, with quotes.
The answer in one line: n_transcripts = 10; objection_count = 3; rep = Alex Franklin (+3 more below)
n_transcripts10
objection_count3
allowed_new13 items — e.g. TT-001, TT-002, TT-003, TT-004
competitor_quotesAwardco=We're also in late talks with Awardco — their rewards catalog looks bigger than yours.; Kudos=How are you different from Kudos? Our CEO used them at her last company.
competitors_prospect_raisedAwardco, Kudos
next_stepagreed=7; total=10; rate=0.7
objectionsobjection=budget; count=4; quote=Honestly, budget is locked until next fiscal year — I can't add a new line item right now.; transcript_id=TT-001 | objection=timing; count=3; quote=This is good, but can we revisit it next quarter? Open enrollment starts in October and we're underwater.; transcript_id=TT-002 | objection=in-house; count=3; quote=We already do recognition with a spreadsheet and quarterly gift cards — why would we change?; transcript_id=TT-004
opener_themetheme=opens with a customer recognition win story (400-person retailer, 18% turnover reduction); count=8; total=10; quote=Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it.; transcript_id=TT-001
repAlex Franklin
rep_raised_competitors_excludedWorkhuman
graded checksobjection_count · next_step_rate · competitors · coaching · quote_used

Model's actual answer

show the model's full answer
First five minutes

- 8/10 calls (80%) open with the same customer-outcome proof point:
  “Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it.”
- Exceptions: Deal-403845 starts with a security-review/pricing agenda; Deal-1E2498 starts with the prospect’s prior request for straight pricing.
- Deal-C61CF7 also adds a Workhuman pricing/catalog comparison by minute 2.

Three most common objections

Prospect objection/competitive lines: 15 total = 6 budget/committee + 3 timing + 3 status quo + 1 no urgency + 2 competitor mentions.

1. Budget/committee constraint — 6/15 lines across 5/10 calls
   - Budget locked: Deal-D348E1, Deal-547B2B, Deal-60C2C2, Deal-84DBA6: 4 lines.
   - Committee dependency: Deal-403845, Deal-84DBA6: 2 lines.
   - Handling: 3/4 budget-locked instances receive an avoided-backfill funding argument:
     “Totally fair. Most teams fund this out of turnover savings — that retailer saved about $210k in avoided backfills, which is how their finance team signed off.”
   - Both committee-related instances end with acknowledgment rather than a proposed path forward.

2. Timing/financial timing — 3/15 lines across 3/10 calls
   - Deal-5408B0, Deal-C61CF7, Deal-D9A12F.
   - Handling: 3/3 acknowledge the concern and propose the same pilot:
     “Makes sense. What if we scope a 90-day pilot with one department so you have internal data before next quarter's planning?”

3. Existing recognition process — 3/15 lines across 3/10 calls
   - Deal-403845, Deal-EDC141, Deal-1E2498.
   - Handling: 3/3 differentiate on automation and analytics:
     “Spreadsheets work until they scale — the difference is automation: milestones fire without HR lifting a finger, and you get analytics on who is being recognized.”

Concrete next-step agreement rate

- Agreed: 7/10 = 70%.
- 7 ÷ 10 × 100 = 70%.
- Agreed aliases: Deal-D348E1, Deal-5408B0, Deal-547B2B, Deal-C61CF7, Deal-60C2C2, Deal-D9A12F, Deal-1E2498.
- Not agreed: 3/10 = 30% — Deal-403845, Deal-EDC141, Deal-84DBA6.
- Agreement evidence:
  “Yes, Thursday at 2pm works — send the invite and I'll bring our HRIS manager.”

Competitors raised by prospects

- Deal-547B2B — Awardco:
  “We're also in late talks with Awardco — their rewards catalog looks bigger than yours.”
- Deal-EDC141 — Kudos:
  “How are you different from Kudos? Our CEO used them at her last company.”

Workhuman is not included because Alex Franklin raised it in Deal-C61CF7; no company-alias field was provided.

Coaching notes

1. Replace the identical opener used in 8/10 calls with a prospect-specific priority. Pricing appears in Deal-C61CF7 and Deal-1E2498; security appears in Deal-403845.
2. Improve the close on resistant calls. Three calls ended without a next step; ask for a dated decision point and agree on the specific follow-up needed after committee review or the prospect’s internal decision.
reporting-analytics 5 tests
q3-weighted-forecast0.90
Plain-English question: Build the Q3 2026 weighted forecast from the open pipeline, and show the math deal by deal.
Prompt — what the model saw
Below are the open deals in the default sales pipeline with a close date between 2026-07-01 and 2026-10-15, pulled from the CRM warehouse on 2026-09-05. Stages: DS1 (discovery) through DS5 (contract). Build the Q3 2026 forecast (quarter = 2026-07-01 to 2026-09-30). Weighting rule: 100 percent of COMMIT plus 35 percent of BEST_CASE. PIPELINE deals count zero. Only deals with a close date inside the quarter count. Report: COMMIT total, BEST_CASE total, weighted forecast, count of deals in each category inside the quarter, the deals you excluded for being outside the quarter (count and total amount), and the top 5 BEST_CASE deals by amount inside the quarter. Show arithmetic. Then, in a final '## Data quality' section of no more than four sentences, name the data quality problems in this extract that would make the forecast unreliable if it were run unattended.
cost $0.0000 · 163s · in 55,541 / out 20,064 tok
✓ commit_total_44729✓ best_case_total_203565✓ weighted_115976.75✓ commit_count_7✓ best_case_count_24✓ excluded_count_32✓ excluded_amount_227575✓ top5_aliases✓ flags_blank_owner✗ flags_missing_why_buys

✓ Correct Answer (ground truth)

Plain-English answer: A stage-weighted Q3 forecast number computed from the open pipeline, with per-stage math that reconciles to the deal data.
The answer in one line: excluded_amount = 227575; best_case_total = 203565; weighted_forecast = 115,976.75; commit_total = 44729; excluded_out_of_quarter = 32; best_case_count = 24; flags = blank owner column + empty why_buys fields must be surfaced (+2 more below)
excluded_amount227,575
best_case_count24
best_case_total203,565
commit_count7
commit_total44,729
excluded_out_of_quarter32
weighted_forecast115,976.75
flagsblank owner column + empty why_buys fields must be surfaced
top5_best_caseDeal-2D7423, Deal-25F752, Deal-E53952, Deal-5EED42, Deal-FA32A0
graded checkscommit_total_44729 · best_case_total_203565 · weighted_115976.75 · commit_count_7 · best_case_count_24 · excluded_count_32 · excluded_amount_227575 · top5_aliases · flags_blank_owner · flags_missing_why_buys

Model's actual answer

show the model's full answer
Q3 2026 forecast — 2026-07-01 through 2026-09-30, inclusive

Amounts are shown in raw extract units because no currency is provided.

| Category | Deals inside quarter | Amount | Forecast contribution |
|---|---:|---:|---:|
| COMMIT | 7 | 44,729 | 44,729 |
| BEST_CASE | 24 | 203,565 | 71,247.75 |
| PIPELINE | 23 | 201,637.40 | 0 |
| Total | 54 | 449,931.40 | 115,976.75 |

COMMIT arithmetic:

11,200 + 9,000 + 9,000 + 6,360 + 5,400 + 2,520 + 1,249 = 44,729

BEST_CASE arithmetic:

38,935 + 24,000 + 19,656 + 16,250 + 11,116 + 10,800 + 10,500 + 9,890 + 9,720 + 9,000 + 7,200 + 3,840 + 3,780 + 3,600 + 3,240 + 3,150 + 3,120 + 3,060 + 2,916 + 2,760 + 2,484 + 2,100 + 1,920 + 528 = 203,565

Weighted forecast:

(1.00 × 44,729) + (0.35 × 203,565) + (0 × 201,637.40)  
= 44,729 + 71,247.75  
= 115,976.75

Excluded outside the quarter

All 32 excluded deals have close dates from 2026-10-01 through 2026-10-15.

| Category | Excluded count | Excluded amount |
|---|---:|---:|
| COMMIT | 1 | 13,770 |
| BEST_CASE | 9 | 28,240 |
| PIPELINE | 22 | 185,565 |
| Total | 32 | 227,575 |

Arithmetic: 1 + 9 + 22 = 32 deals; 13,770 + 28,240 + 185,565 = 227,575.

Excluded aliases:

- COMMIT: Deal-D348E1
- BEST_CASE: Deal-C61CF7, Deal-48B656, Deal-901332, Deal-47AE31, Deal-15D24F, Deal-ED725A, Deal-8AD4A5, Deal-5FDCE4, Deal-F5A622
- PIPELINE: Deal-E51FB7, Deal-B936FE, Deal-D9A12F, Deal-4062CF, Deal-293AF3, Deal-034D49, Deal-E0ADD8, Deal-9F2E43, Deal-FCBE5B, Deal-712010, Deal-6691E0, Deal-600CD9, Deal-A92065, Deal-1D532E, Deal-E531A6, Deal-D1E6C2, Deal-D9E112, Deal-5AD94B, Deal-766C74, Deal-D7E999, Deal-ED13B0, Deal-7FA0C3

Top 5 BEST_CASE deals inside the quarter

| Rank | Deal alias | Amount |
|---:|---|---:|
| 1 | Deal-2D7423 | 38,935 |
| 2 | Deal-25F752 | 24,000 |
| 3 | Deal-E53952 | 19,656 |
| 4 | Deal-5EED42 | 16,250 |
| 5 | Deal-FA32A0 | 11,116 |

## Data quality

The extract has no currency field, so the totals are not independently verifiable as currency-comparable. It is a 2026-09-05 point-in-time snapshot with no last-updated, category-change, or explicit open/closed status field, so recency and status cannot be audited without a re-pull. Four rows labeled open were already past their close date at pull time—Deal-333EBB, Deal-57FF13, Deal-31AD2C, and Deal-7A2454—indicating possible stale close-state handling; all four are PIPELINE and contribute zero. Owner is blank in 85 of 86 rows, and Deal-42326B has the extract’s only fractional amount (2,480.40), so completeness and amount-granularity checks are not fully automatable.
aha-moment-2x20.83
Plain-English question: Do companies that hit 5+ givers and one redemption in month one actually retain better at 24 months?
Prompt — what the model saw
Test the activation hypothesis on the company cohort below: companies with 5+ unique givers (m1_users >= 5) AND 1+ successful redemption (m1_redemptions >= 1) in their first calendar month retain better at 24 months than those with one signal or neither. Every company in this extract is already 25+ months old, so current_status = 'active' means retained at 24 months. Report the full 2x2 (cohort size and 24-month retention rate for: both signals, givers-only, redemption-only, neither), state how many companies were excluded from the denominator and why (if none, say so), name the single signal with the largest retention lift, and state what this does and does not prove.
cost $0.0000 · 195s · in 58,854 / out 17,191 tok
✓ both_66✓ givers_47✓ redemp_31✓ neither_40✓ excluded_0✗ largest_lift_givers

✓ Correct Answer (ground truth)

Plain-English answer: The four cohort cells are graded against computed 24-month retention; the numbers come from the scorer's cohort table, not opinion.
The answer in one line: both_5plus_givers_and_redemption = 66.0% (31 of 47)
both_5plus_givers_and_redemption66.0% (31 of 47)
exclusionsnone — claim bounded as correlation, not causation
givers_only46.9% (23 of 49)
largest_single_signal_liftunique givers
neither40.0% (38 of 95)
redemption_only31.0% (9 of 29)
graded checksboth_66 · givers_47 · redemp_31 · neither_40 · excluded_0 · largest_lift_givers

Model's actual answer

show the model's full answer
Definitions applied:
- Givers signal G: m1_users ≥ 5
- Redemption signal R: m1_redemptions ≥ 1
- Retained: current_status = 'active'; 'cancelled' and 'non_renewing' are not retained

24-MONTH RETENTION

Cohort                Cohort size   Retained   Arithmetic
Both signals           47           31         31/47 = 65.96%
Givers-only            49           23         23/49 = 46.94%
Redemption-only        29            9          9/29 = 31.03%
Neither                95           38         38/95 = 40.00%
Total                 220          101        101/220 = 45.91%

Cell-size check: 47 + 49 + 29 + 95 = 220.
Retained check: 31 + 23 + 9 + 38 = 101.

Excluded from denominator: 0 companies. All 220 rows were included. The three non_renewing companies remain in the denominator and count as not retained.

HYPOTHESIS TEST

Both signals versus:
- Givers-only: 65.96% − 46.94% = +19.02 percentage points
- Redemption-only: 65.96% − 31.03% = +34.92 points
- Neither: 65.96% − 40.00% = +25.96 points

The activation hypothesis is descriptively supported in this extract: companies with both signals had the highest observed 24-month retention.

LARGEST SINGLE-SIGNAL LIFT

Givers (m1_users ≥ 5).

Using the one-signal cells versus neither:
- Givers-only: 46.94% − 40.00% = +6.94 points
- Redemption-only: 31.03% − 40.00% = −8.97 points

The conclusion is also the same using marginal comparisons:
- Givers: 54/96 = 56.25% versus 47/124 = 37.90%, a +18.35-point lift
- Redemptions: 40/76 = 52.63% versus 61/144 = 42.36%, a +10.27-point lift

WHAT THIS PROVES

Within this 220-company extract, having both first-month signals is associated with better observed 24-month retention, and the givers signal has the larger standalone association.

WHAT THIS DOES NOT PROVE

It does not prove that causing companies to reach these thresholds will cause higher retention, that the differences will generalize beyond this extract, or that the 5-user and one-redemption thresholds are optimal. It is observational and does not isolate the signals from other differences between companies.
arr-rollforward-reconciliation0.80
Plain-English question: Does CRM ARR match billing ARR — and if not, exactly where is the variance?
Prompt — what the model saw
Reconcile ARR between CRM company records and billing subscriptions as of 2026-09-05. Billing ARR is MRR x 12. Report both totals, the variance, and a decomposition into buckets (status mismatch, rounding, missing records, other) that sums exactly to the variance. List each mismatched account with a suggested owner. Business rule: any subscription with a term other than 12 months must have cf_agreement_end_date populated; list violations.
cost $0.0000 · 133s · in 33,736 / out 15,458 tok
✓ billing_total✓ crm_total✗ variance✓ buckets_sum✓ violations_listed

✓ Correct Answer (ground truth)

Plain-English answer: Billing $604,739.28 vs CRM $603,581.76: a $1,157.52 variance decomposed into status-mismatch, rounding, and missing-record buckets.
The answer in one line: n_companies = 40; n_mismatched_accounts = 7; as_of = 2026-09-05 (+4 more below)
n_companies40
n_mismatched_accounts7
as_of2026-09-05
billing_arr_total604739.28
bucket_sum_equals_variance1
buckets{"status_mismatch": {"amount": "-13158.48", "accounts": [{"company_alias": "C-0C8323BF", "billing_arr": "0.00", "hubspot_arr": "4905.24", "diff": "-4905.24", "suggested_owner": "Dana Mercer"}, {"compa
crm_arr_total603581.76
ruleterm_months != 12 requires cf_agreement_end_date
variance1157.52
variance_definitionbilling_arr_total - crm_arr_total (billing ARR = sum of MRR x 12 over status=active subscriptions)
violationssubscription_id=SUB-0002; company_alias=C-1794A52C; term_months=24 | subscription_id=SUB-0019; company_alias=C-22170CA1; term_months=36
graded checksbilling_total · crm_total · variance · buckets_sum · violations_listed

Model's actual answer

show the model's full answer
As of 2026-09-05

Currency is not specified; amounts are shown in source currency units. All 39 subscriptions and all 39 CRM records are included, including cancelled subscriptions.

TOTALS AND ARITHMETIC

Billing ARR:
- Active MRR: 50,394.94 × 12 = 604,739.28
- Cancelled MRR: 1,096.54 × 12 = 13,158.48
- Total MRR: 50,394.94 + 1,096.54 = 51,491.48
- Billing ARR: 51,491.48 × 12 = 617,897.76

CRM ARR:
- Sum of CRM records:
  132,638.64 + 127,977.00 + 169,075.56 + 173,890.56
  = 603,581.76

Variance, Billing − CRM:
- 617,897.76 − 603,581.76 = +14,316.00
- Billing ARR is higher by 14,316.00.

VARIANCE DECOMPOSITION

| Bucket | Billing − CRM | Arithmetic |
|---|---:|---|
| Status mismatch | 0.00 | No quantifiable amount; both cancelled subscriptions have equal CRM ARR |
| Rounding | -36.00 | -16.00 + -20.00 |
| Missing records | +11,952.00 | +28,449.24 billing-only − 16,497.24 CRM-only |
| Other | +2,400.00 | 26,796.00 billing − 24,396.00 CRM |
| Total | +14,316.00 | 0.00 - 36.00 + 11,952.00 + 2,400.00 = 14,316.00 |

“Rounding” and “other” are reconciliation classifications; the supplied data contains no discrepancy reason codes. The CRM file also has no status field, so status agreement itself cannot be verified.

MISMATCHED ACCOUNTS

No named owners were supplied. Suggested functional owner for each mismatch: Revenue Operations.

| company_alias | subscription_id | Billing ARR | CRM ARR | Variance: Billing − CRM | Bucket | Suggested owner |
|---|---|---:|---:|---:|---|---|
| C-0D66DF9E | SUB-0005 | 23,184.00 | 23,200.00 | -16.00 | Rounding | Revenue Operations |
| C-0F7269D7 | SUB-0006 | 26,796.00 | 24,396.00 | +2,400.00 | Other | Revenue Operations |
| C-14D70CE0 | SUB-0008 | 18,180.00 | 18,200.00 | -20.00 | Rounding | Revenue Operations |
| C-21629AA4 | SUB-0004 | 28,449.24 | Not present | +28,449.24 | Missing records | Revenue Operations |
| C-0D5BBE3A | No matching subscription | Not present | 16,497.24 | -16,497.24 | Missing records | Revenue Operations |

Account-level check:
-16.00 + 2,400.00 - 20.00 + 28,449.24 - 16,497.24 = 14,316.00.

TERM-DATE VIOLATIONS

| subscription_id | company_alias | term_months | status | cf_agreement_end_date |
|---|---|---:|---|---|
| SUB-0002 | C-1794A52C | 24 | active | Blank |
| SUB-0019 | C-22170CA1 | 36 | active | Blank |

Total violations: 2.
kvm-month-over-monthPASS 1.00
Plain-English question: How did the four core KVMs move in August versus July?
Prompt — what the model saw
Report the four core KVMs for 2026-08 against 2026-07: giving rate, redemptions per user, 1:1 meetings engagement, pulse check engagement. For each: value, prior value, absolute and relative change, direction. Name the KVM with the largest relative move and the segment (plan_tier or size_band) driving it if the data supports that; otherwise say the data does not. Under 200 words plus one table.
cost $0.0000 · 53s · in 8,662 / out 6,230 tok
✓ four_kvms✓ biggest_mover_named✓ direction_correct✓ segment_named✓ word_cap

✓ Correct Answer (ground truth)

Plain-English answer: Giving rate 60.27% (up 0.07%), redemptions per user 1.7302 (up 0.01%), with 1:1 meetings and pulse engagement each up fractionally — direction reported for all four.
The answer in one line: n_company_rows_per_month = 30; aggregation = simple mean of company-level values per month
n_company_rows_per_month30
aggregationsimple mean of company-level values per month
biggest_moverpulse_engagement
driving_segmentsize_band=enterprise; 2026-07=0.55; 2026-08=0.2743; rel_change_pct=-50.13
kvms{"giving_rate": {"2026-07": 0.6023, "2026-08": 0.6027, "abs_change": 0.0004, "rel_change_pct": 0.07, "direction": "up"}, "redemptions_per_user": {"2026-07": 1.73, "2026-08": 1.7302, "abs_change": 0.00
segment_detail{"smb": {"2026-07": 0.6588, "2026-08": 0.6573, "rel_change_pct": -0.23}, "mid_market": {"2026-07": 0.593, "2026-08": 0.5942, "rel_change_pct": 0.2}, "enterprise": {"2026-07": 0.55, "2026-08": 0.2743,
graded checksfour_kvms · biggest_mover_named · direction_correct · segment_named · word_cap

Model's actual answer

show the model's full answer
Method: Unweighted arithmetic mean across 30 matched company aliases; each company has both months. No user/company weights were provided. Aug = Σ values ÷ 30; prior = Σ values ÷ 30; relative change = (Aug − prior) ÷ prior. Absolute changes are in raw units.

| Core KVM | 2026-08 value | 2026-07 prior | Absolute change | Relative change | Direction |
|---|---:|---:|---:|---:|---|
| Giving rate | 18.081400 ÷ 30 = 0.602713 | 18.068900 ÷ 30 = 0.602297 | +0.000417 | +0.0692% | Up |
| Redemptions per user | 51.904900 ÷ 30 = 1.730163 | 51.899500 ÷ 30 = 1.729983 | +0.000180 | +0.0104% | Up |
| 1:1 meetings engagement | 13.415300 ÷ 30 = 0.447177 | 13.406600 ÷ 30 = 0.446887 | +0.000290 | +0.0649% | Up |
| Pulse check engagement | 15.258300 ÷ 30 = 0.508610 | 18.017600 ÷ 30 = 0.600587 | −0.091977 | −15.3145% | Down |

Largest relative move: pulse check engagement, down 15.31%.

Driver: size_band=enterprise. Its mean fell from 0.549980 to 0.274280: change = −0.275700; relative change = −0.275700 ÷ 0.549980 = −50.13%. SMB declined 0.22%, while mid_market increased 0.21%, so enterprise drove the portfolio decline. A plan_tier comparison is unsupported because only tier_three is present.
rewards-ytd-sectionPASS 1.00
Plain-English question: Produce the redemption section of the weekly rewards report through August.
Prompt — what the model saw
Produce the redemption section of the weekly rewards report through the last completed month (2026-08): redemption count, spend, unique redeemers, redemptions per redeemer, provider mix as percent of spend, top 5 countries by redemptions. Say which month is the last completed month. Provider shares must sum to 100.
cost $0.0000 · 236s · in 95,389 / out 23,436 tok
✓ count✓ spend✓ unique_redeemers✓ month_named✓ provider_mix✓ top_country

✓ Correct Answer (ground truth)

Plain-English answer: 378 redemptions, $27,846 spend, 235 unique redeemers at 1.61 each, custom rewards at 39% of spend, US leading all countries.
The answer in one line: redemption_count = 378; unique_redeemers = 235; redemptions_per_redeemer = 1.61; total_spend_usd = 27846.00
redemption_count378
redemptions_per_redeemer1.61
unique_redeemers235
provider_mix_pct_of_spendTangoCard=11.6; Tremendous=30.54; custom=39.05; Snappy=18.81
top5_countries_by_countcountry=US; count=244 | country=CA; count=24 | country=AU; count=21 | country=GB; count=17 | country=NL; count=17
total_spend_usd27846.00
graded checkscount · spend · unique_redeemers · month_named · provider_mix · top_country

Model's actual answer

show the model's full answer
REDEMPTION SECTION — YTD THROUGH 2026-08

Last completed month: 2026-08 (August 2026)  
Period: 2026-01-01 through 2026-08-31  
Source: provided redemptions_ytd.csv

Redemption count: 377  
Spend: $27,836.00  
Unique redeemers: 234 distinct user_key values  
Redemptions per redeemer: 377 ÷ 234 = 1.6111 → 1.61

Provider mix — percent of spend

Provider       Redemptions  Spend        Calculation                         Share
custom         37           $10,873.00   $10,873 ÷ $27,836 = 39.0609%         39.1%
Tremendous     191          $8,495.00    $8,495 ÷ $27,836 = 30.5180%          30.5%
Snappy         59           $5,238.00    $5,238 ÷ $27,836 = 18.8174%          18.8%
TangoCard      90           $3,230.00    $3,230 ÷ $27,836 = 11.6037%          11.6%
TOTAL          377          $27,836.00                                      100.0%

Top 5 countries by redemptions

Rank  Country  Redemptions
1     US       243
2     CA       24
3     AU       21
4=    GB       17
4=    NL       17

GB and NL are tied at 17 redemptions.
customer-success 5 tests
churn-save-eligibilityPASS 1.00
Plain-English question: Which at-risk accounts actually qualify for a churn-save offer, and which play fits each?
Prompt — what the model saw
Which at-risk accounts qualify for a churn-save offer under the documented eligibility rules, what amount is at stake per account and in total, and which play fits each (usage revival, executive touch, commercial concession)? Cite the signal that justifies each play. List accounts that look at risk but do not qualify and why.
cost $0.0000 · 153s · in 36,719 / out 17,420 tok
✓ eligible_set✓ total_at_stake✓ plays_cited✓ noneligible_named✓ rules_applied

✓ Correct Answer (ground truth)

Plain-English answer: Only accounts meeting all three rules qualify — e.g. C-0F6C0F34 at $49,707 (executive touch) and C-0B827671 at $25,365 (usage revival) — with amount and play per account.
The answer in one line: n_accounts = 30; snapshot_date = 2026-09-05
n_accounts30
eligibleaccount_alias=C-0F6C0F34; amount_at_stake=49707.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0B827671; amount_at_stake=25365.00; play=usage revival; justifying_signal=usage_trend_3m=declining over the last 3 months | account_alias=C-0B360C78; amount_at_stake=35748.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 75% - risk is commercial, not adoption | account_alias=C-0B0F1BAB; amount_at_stake=5494.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0CA21961; amount_at_stake=16829.00; play=usage revival; justifying_signal=seat utilization 26% is below 50% | account_alias=C-0E9C27D1; amount_at_stake=41235.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 85% - risk is commercial, not adoption … (+2 more)
non_eligible_at_riskaccount_alias=C-0BC71BDD; health_score=55 | account_alias=C-0BA71F12; health_score=52 | account_alias=C-0F6694C3; health_score=43 | account_alias=C-0BE96399; health_score=54 | account_alias=C-0F876796; health_score=47 | account_alias=C-0FCCD2DF; health_score=43 … (+1 more)
ruleshealth_score < 60, churn_save_eligible_amount > 0, renewal within 120 days of snapshot
snapshot_date2026-09-05
total_amount_at_stake224601.00
graded checkseligible_set · total_at_stake · plays_cited · noneligible_named · rules_applied

Model's actual answer

show the model's full answer
Eligibility applied:
- R1: health_score < 60
- R2: churn_save_eligible_amount > 0
- R3: renewal_date from 2026-09-05 through 2027-01-03
  - 2026-09-05 + 120 days = 2027-01-03

The files document eligibility only, not play-selection rules. Play recommendations below use only the supplied signals; no play is forced where the data does not support one.

QUALIFYING ACCOUNTS — 8

Account          Days to renewal   Amount at stake   Best-fit play             Supporting signal
C-0B0F1BAB       18                5,494.00         Executive touch           champion_active=false
C-0E9C27D1       19                41,235.00        Play not determinable      Usage=flat; champion_active=true; 134/157 seats used=85.4%
C-0F6C0F34       28                49,707.00        Executive touch           champion_active=false
C-0B360C78       53                35,748.00        Play not determinable      Usage=growing; champion_active=true; 246/327 seats used=75.2%
C-0D3278C7       68                17,602.00        Usage revival             usage_trend_3m=declining
C-0B827671       70                25,365.00        Usage revival             usage_trend_3m=declining
C-0CEF69FD       77                32,621.00        Executive touch           champion_active=false
C-0CA21961       114               16,829.00        Commercial concession    84/325 seats used=25.8%; 325−84=241 unused seats

Total amount at stake:
49,707.00
+ 25,365.00
+ 35,748.00
+ 5,494.00
+ 16,829.00
+ 41,235.00
+ 32,621.00
+ 17,602.00
= 224,601.00

No currency field is provided, so amounts are shown exactly as supplied.

AT-RISK BUT NOT ELIGIBLE — 7

C-0BC71BDD: R1 passes (health 55); fails R2 because eligible amount=0.00. Renewal is 52 days away, so R3 passes.

C-0BA71F12: R1 passes (health 52); R2 passes (6,824.00); fails R3 because renewal is 218 days after the snapshot, exceeding 120 days.

C-0F6694C3: R1 passes (health 43); fails R2 because eligible amount=0.00 and fails R3 because renewal is 197 days away.

C-0BE96399: R1 passes (health 54); fails R2 because eligible amount=0.00. Renewal is 54 days away, so R3 passes.

C-0F876796: R1 passes (health 47); R2 passes (19,958.00); fails R3 because renewal is 154 days after the snapshot.

C-0FCCD2DF: R1 passes (health 43); fails R2 because eligible amount=0.00 and fails R3 because renewal is 230 days away.

C-10A56B0F: R1 passes (health 54); fails R2 because eligible amount=0.00. Renewal is 98 days away, so R3 passes.
expansion-kitPASS 1.00
Plain-English question: Build the expansion kit for the customer who just replied to our outreach.
Prompt — what the model saw
A customer replied to outreach (reply included). Build the expansion kit: seat coverage (licensed over headcount), two lines on usage health, headroom in seats and ARR at the current per-seat rate, who replied and whether they can buy, the right buyer from our contacts, and a reply email under 150 words that moves toward the right person without being pushy and cites one real usage fact.
cost $0.0000 · 49s · in 43,592 / out 4,378 tok
✓ coverage_pct✓ headroom_arr✓ right_buyer✓ replier_not_buyer✓ usage_fact✓ email_under_150

✓ Correct Answer (ground truth)

Plain-English answer: 37.5% seat coverage, 250 headroom seats worth $15,000 ARR at $60 a seat, usage up 43.2% in six months — and the replier cannot buy; Dana R. (VP People) is the right buyer.
The answer in one line: headroom_seats = 250; coverage_pct = 37.5; company_alias = C-0DDFC9A7 (+3 more below)
headroom_seats250
coverage_pct37.5
company_aliasC-0DDFC9A7
headroom_arr15000.00
per_seat_annual_rate60.00
replierMaria S.
replier_can_buy0
right_buyername=Dana R.; title=VP People
usage_factsactive_users_2026_08=126; active_users_2026_03=88; growth_pct_6m=43.2; licensed_seat_utilization_pct=84
usage_health_linesMonthly active users grew from 88 (2026-03) to 126 (2026-08), +43.2% over 6 months., Last month 126 of 150 licensed seats were active (84.0% utilization), approaching the seat cap.
graded checkscoverage_pct · headroom_arr · right_buyer · replier_not_buyer · usage_fact · email_under_150

Model's actual answer

show the model's full answer
C-0DDFC9A7

Seat coverage
- 150 licensed ÷ 400 headcount = 37.5% coverage.
- Coverage shortfall: 400 − 150 = 250 seats.

Usage health
- Users increased every month: 88 → 95 → 102 → 110 → 118 → 126. March-to-August growth: 126 − 88 = 38 users, or 38 ÷ 88 = 43.2%.
- August usage: 126 ÷ 150 = 84.0% of licensed capacity, leaving 150 − 126 = 24 seats unused.

Headroom and value
- Current ARR per licensed seat: $9,000 ÷ 150 = $60 per seat.
- Full-headcount headroom: 250 seats × $60 = $15,000 in additional ARR.
- Current unused-license capacity: 24 seats × $60 = $1,440 in ARR represented by already-licensed seats.

Contacts
- Replied: Maria S., People Operations Coordinator; last engaged 2026-09-02.
- Can Maria buy? No. She explicitly says purchasing decisions are not hers and that budget and seat expansion sit with Dana R.
- Right buyer: Dana R., VP People; last engaged 2026-05-18. Maria also offered to make the introduction.

Reply email

Subject: RE: Growing your team's recognition program

Hi Maria,

Thanks for the thoughtful reply and for confirming where budget and seat decisions sit. We’re glad the team is getting value from Bonusly. One concrete usage data point: users increased from 88 in March to 126 in August.

Would you be comfortable introducing me to Dana? I’d like to understand what seat coverage looks like as your team grows. If an introduction is easier, I can send a short note she can forward.

Best,
Cole
onboarding-statusPASS 1.00
Plain-English question: What is done and what is missing in this customer's onboarding before the mid-onboarding call?
Prompt — what the model saw
Prep the CSM for the mid-onboarding call: what is complete (integrations connected, allowance set, admins added, first recognition given, first redemption), what is not, early engagement signals, and the three things to cover on the call. Mark nothing complete without a data field that shows it.
cost $0.0000 · 47s · in 29,345 / out 3,855 tok
✓ complete_items✓ incomplete_items✓ three_items✓ grounded

✓ Correct Answer (ground truth)

Plain-English answer: 25 days live: Slack connected, allowance set, admins added — but the HRIS integration is not connected, with early engagement signals and three call priorities.
The answer in one line: days_live = 25; company_alias = C-0D284E42
days_live25
call_itemsConnect the HRIS integration (integration_hris is empty) to automate user provisioning and offboarding., Get to first redemption (first_redemption_at is empty): walk admins through the rewards catalog and confirm allowance funding., Build the giving habit: daily active givers averaged 4.3 in week 1 vs 13.0 last week - agree on manager nudges and a recognition cadence.
checklist{"slack_integration_connected": {"complete": true, "field": "integration_slack", "value": "2026-08-12"}, "hris_integration_connected": {"complete": false, "field": "integration_hris", "value": ""}, "a
company_aliasC-0D284E42
engagement_signalsavg_active_givers_week1=4.3; avg_active_givers_last_week=13; latest_active_givers=15; days_with_activity=25
graded checkscomplete_items · incomplete_items · three_items · grounded

Model's actual answer

show the model's full answer
C-0D284E42

Complete
- Slack integration: `integration_slack=2026-08-12` — connected 1 day after signup.
- Allowance: `allowance_set=2026-08-13` — set 2 days after signup.
- Admins: `admins_added=2` — completion recorded; timing unavailable.
- First recognition: `first_recognition_at=2026-08-15 14:22` — given 4 calendar days after signup.

Not complete in the provided data
- HRIS integration: `integration_hris` is blank; no connection date.
- First redemption: `first_redemption_at` is blank; no redemption recorded.

Integration arithmetic: 1 of 2 integration fields is populated = 1 ÷ 2 = 50%.

Early engagement signals
- Active givers increased from 3 on 2026-08-11 to 15 on 2026-09-04:
  - Net increase: 15 − 3 = 12.
  - Ending level: 15 ÷ 3 = 5× the signup-day level.
  - Percentage increase: 12 ÷ 3 = 400%.
- By the first-recognition date, 2026-08-15, active givers were 5: 5 − 3 = 2 above signup day.
- Usage is provided for 25 dates, 2026-08-11 through 2026-09-04 inclusive; no later usage is available.
- Total eligible users were not provided, so participation rate cannot be calculated.

Three things to cover on the call
1. Complete the HRIS integration: confirm the target system, owner, blockers, and target connection date.
2. Drive the first redemption: confirm the recipient, reward, internal owner, and target redemption date.
3. Scale participation beyond 15 active givers: identify the next cohort to engage and confirm how the 2 added admins will coordinate the rollout.
renewal-risk-conflicting-datesPASS 1.00
Plain-English question: Two systems disagree on renewal dates — which do you trust per account, and who is actually renewing inside 90 days?
Prompt — what the model saw
Produce a 90-day renewal risk brief. Two systems hold renewal dates and disagree on some accounts; decide which to trust per account and say why (multi-year contracts are known to be wrong in ChurnZero). For every renewal: company, CSM, ARR, date used, seat utilization, 3-month usage trend, risk rating with one sentence of evidence. Flag every disagreement. Close with total ARR renewing and ARR at risk.
cost $0.0000 · 233s · in 36,493 / out 27,847 tok
✓ total_renewing✓ arr_at_risk✓ disagreements_flagged✓ trust_rule

✓ Correct Answer (ground truth)

Plain-English answer: Chargebee wins per account (multi-year contracts are known wrong in ChurnZero), with each trusted date, the reason, and a 90-day-window flag.
The answer in one line: n_accounts = 20; n_disagreements = 5; snapshot_date = 2026-09-05 (+2 more below)
n_accounts20
n_disagreements5
accounts20 items — e.g. account_alias=C-0B144C78; csm=Cole Ingram; arr=30899.00; trusted_renewal_date=2026-11-02; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.4; usage_3m_ratio=1.03; risk=low; evidence=3-month usage ratio 1.03 (last3 avg 103 vs prior3 100), seat utilization 75% | account_alias=C-0B20DB64; csm=Dana Mercer; arr=21770.00; trusted_renewal_date=2026-10-07; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=56.6; usage_3m_ratio=1; risk=medium; evidence=3-month usage ratio 1.00 (last3 avg 295 vs prior3 295), seat utilization 57% | account_alias=C-0B344485; csm=Elena Sinclair; arr=64384.00; trusted_renewal_date=2026-11-16; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=78; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 241 vs prior3 231), seat utilization 78% | account_alias=C-0B369871; csm=Dana Mercer; arr=85128.00; trusted_renewal_date=2026-10-29; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.1; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 330 vs prior3 316), seat utilization 75%
arr_at_risk_90d359409.00
disagreementsaccount_alias=C-0B7D2C30; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-15; is_multi_year=True; trusted=2026-09-15 | account_alias=C-0BBE3E60; cz_renewal_date=2027-09-26; cb_renewal_date=2026-09-26; is_multi_year=True; trusted=2026-09-26 | account_alias=C-0BCDB8C2; cz_renewal_date=2027-09-18; cb_renewal_date=2026-09-18; is_multi_year=True; trusted=2026-09-18 | account_alias=C-0D2AB865; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-22; is_multi_year=True; trusted=2026-09-22 | account_alias=C-0F5D2323; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-29; is_multi_year=True; trusted=2026-09-29
snapshot_date2026-09-05
total_arr_renewing_90d1048715.00
trust_rulemulti-year contracts: Chargebee is authoritative (ChurnZero known wrong); otherwise systems agree or Chargebee wins
window2026-09-05 to 2026-12-04
graded checkstotal_renewing · arr_at_risk · disagreements_flagged · trust_rule

Model's actual answer

show the model's full answer
90-DAY RENEWAL RISK BRIEF

Scope and method
- No report/as-of date is supplied. The 90-day window is therefore assumed to be 2026-09-01–2026-11-29, beginning after the latest usage month, 2026-08. No renewals are supplied for November 25–29.
- No company names or currency are provided. Company uses account_alias exactly; ARR is shown in the file’s unspecified monetary units.
- Utilization = seats_used ÷ seats. Three-month trend = 2026-06→07→08 active_users; change = (August − June) ÷ June.
- Risk heuristic: High = utilization <40% or usage change ≤−10%; Low = utilization ≥70% and usage change >−5%; Medium = all other cases. At risk means High only.
- Date rule: Chargebee is trusted when is_multi_year=true because the supplied instruction says ChurnZero is wrong for multi-year contracts; both systems are used when dates and terms agree.

Company | CSM | ARR | Date used / source | Seat utilization | Jun→Jul→Aug; change | Risk and evidence
---|---|---:|---|---|---|---
C-0B7D2C30 | Dana Mercer | 65,901.00 | 2026-09-15; Chargebee, 36m | 274/476=57.6% | 97→94→84; (84−97)/97=−13.4% | High — Seat utilization was 57.6% and active users fell 13.4%.
C-0D2AB865 | Elena Sinclair | 38,022.00 | 2026-09-22; Chargebee, 24m | 250/407=61.4% | 125→117→109; (109−125)/125=−12.8% | High — Seat utilization was 61.4% and active users fell 12.8%.
C-0BBE3E60 | Dana Mercer | 30,993.00 | 2026-09-26; Chargebee, 24m | 74/114=64.9% | 39→35→33; (33−39)/39=−15.4% | High — Seat utilization was 64.9% and active users fell 15.4%.
C-0F5D2323 | Cole Ingram | 90,647.00 | 2026-09-29; Chargebee, 24m | 111/390=28.5% | 20→21→18; (18−20)/20=−10.0% | High — Seat utilization was 28.5% and active users fell 10.0%.
C-0EC6999D | Elena Sinclair | 79,419.00 | 2026-10-03; both agree | 31/112=27.7% | 17→16→15; (15−17)/17=−11.8% | High — Seat utilization was 27.7% and active users fell 11.8%.
C-0B20DB64 | Dana Mercer | 21,770.00 | 2026-10-07; both agree | 214/378=56.6% | 294→298→294; (294−294)/294=0.0% | Medium — Seat utilization was 56.6% and active users were flat.
C-0BBC4E7A | Cole Ingram | 56,374.00 | 2026-10-10; both agree | 228/337=67.7% | 142→141→139; (139−142)/142=−2.1% | Medium — Seat utilization was 67.7% and active users fell 2.1%.
C-0FD551AB | Elena Sinclair | 48,815.00 | 2026-10-14; both agree | 210/376=55.9% | 123→122→126; (126−123)/123=+2.4% | Medium — Seat utilization was 55.9% and active users rose 2.4%.
C-0F9F8F13 | Dana Mercer | 46,230.00 | 2026-10-18; both agree | 199/352=56.5% | 185→185→182; (182−185)/185=−1.6% | Medium — Seat utilization was 56.5% and active users fell 1.6%.
C-0BC34584 | Cole Ingram | 16,740.00 | 2026-10-22; both agree | 327/494=66.2% | 104→104→106; (106−104)/104=+1.9% | Medium — Seat utilization was 66.2% and active users rose 1.9%.
C-0B7A7546 | Elena Sinclair | 35,062.00 | 2026-10-25; both agree | 182/205=88.8% | 64→65→63; (63−64)/64=−1.6% | Low — Seat utilization was 88.8% and active users fell only 1.6%.
C-0B369871 | Dana Mercer | 85,128.00 | 2026-10-29; both agree | 317/422=75.1% | 326→330→333; (333−326)/326=+2.1% | Low — Seat utilization was 75.1% and active users rose 2.1%.
C-0B144C78 | Cole Ingram | 30,899.00 | 2026-11-02; both agree | 169/224=75.4% | 101→101→106; (106−101)/101=+5.0% | Low — Seat utilization was 75.4% and active users rose 5.0%.
C-0FC4DBB8 | Elena Sinclair | 94,732.00 | 2026-11-05; both agree | 356/464=76.7% | 189→191→193; (193−189)/189=+2.1% | Low — Seat utilization was 76.7% and active users rose 2.1%.
C-0D5BBE3A | Dana Mercer | 39,740.00 | 2026-11-09; both agree | 85/102=83.3% | 88→90→91; (91−88)/88=+3.4% | Low — Seat utilization was 83.3% and active users rose 3.4%.
C-0FB9D5AF | Cole Ingram | 63,158.00 | 2026-11-13; both agree | 144/199=72.4% | 173→173→176; (176−173)/173=+1.7% | Low — Seat utilization was 72.4% and active users rose 1.7%.
C-0B344485 | Elena Sinclair | 64,384.00 | 2026-11-16; both agree | 224/287=78.0% | 238→240→244; (244−238)/238=+2.5% | Low — Seat utilization was 78.0% and active users rose 2.5%.
C-0CB2C1B4 | Dana Mercer | 40,628.00 | 2026-11-20; both agree | 386/473=81.6% | 47→48→49; (49−47)/47=+4.3% | Low — Seat utilization was 81.6% and active users rose 4.3%.
C-22170CA1 | Cole Ingram | 45,646.00 | 2026-11-24; both agree | 251/294=85.4% | 143→148→146; (146−143)/143=+2.1% | Low — Seat utilization was 85.4% and active users rose 2.1%.

Date disagreements — all four flagged
- C-0B7D2C30: ChurnZero=2026-09-10; Chargebee=2026-09-15. Used Chargebee because it records a 36-month, is_multi_year=true contract.
- C-0D2AB865: ChurnZero=2026-09-10; Chargebee=2026-09-22. Used Chargebee because it records a 24-month, is_multi_year=true contract.
- C-0BBE3E60: ChurnZero=2027-09-26; Chargebee=2026-09-26. Used Chargebee because it records a 24-month, is_multi_year=true contract.
- C-0F5D2323: ChurnZero=2026-09-10; Chargebee=2026-09-29. Used Chargebee because it records a 24-month, is_multi_year=true contract.

Out-of-window record
- C-0BCDB8C2 | Cole Ingram | ARR 54,427.00 | 2027-09-18; both systems agree | 232/424=54.7% | 127→118→110; (110−127)/127=−13.4% | High — Seat utilization was 54.7% and active users fell 13.4%; excluded from 90-day totals because the trusted date is in 2027.

TOTAL ARR RENEWING — 19 in-window renewals
- High: 304,982.00
- Medium: 189,929.00
- Low: 499,377.00
- Reconciliation: 304,982.00 + 189,929.00 + 499,377.00 = 994,288.00.
- Data check: 994,288.00 + 54,427.00 excluded 2027 renewal = 1,048,715.00, equal to the sum of all 20 supplied records.

ARR AT RISK — High tier only
- 65,901.00 + 38,022.00 + 30,993.00 + 90,647.00 + 79,419.00 = 304,982.00.
- 304,982.00 ÷ 994,288.00 × 100 = 30.7% of renewing ARR.

HIGH + MEDIUM WATCHLIST
- 304,982.00 + 189,929.00 = 494,911.00.
- 494,911.00 ÷ 994,288.00 × 100 = 49.8% of renewing ARR.
ticket-theme-synthesis0.80
Plain-English question: What are the real themes in this quarter's support tickets, ignoring the existing tags?
Prompt — what the model saw
Synthesize the quarter's support tickets into 4 to 7 themes. Do not trust existing tags; read the text. Per theme: name, count, share, distinct accounts, ARR affected, two ticket ids, one-line recommendation. Separate broad patterns from single-account noise and rank by ARR exposure, not volume.
cost $0.0000 · 109s · in 31,641 / out 14,374 tok
✗ theme_count_range✓ counts_verifiable✓ single_account_flagged✓ arr_ranked✓ ticket_ids

✓ Correct Answer (ground truth)

Plain-English answer: 80 tickets in 4-7 themes — points-not-posting leads (20 tickets, 25%, $31,100 ARR affected) — each with counts, distinct accounts, ARR, and a recommendation.
The answer in one line: total_tickets = 80; quarter_window = 2026-06-01 through 2026-08-31
total_tickets80
noteexisting tags are noisy/unreliable; classify by text
quarter_window2026-06-01 through 2026-08-31
ranking_by_arr_exposurehris_provisioning, redemption_failures, billing_dispute, points_not_posting, slack_sync
sample_ticket_ids80 items — e.g. IC-460004, IC-460016, IC-460001, IC-460041
single_account_noise_themebilling_dispute
themes{"points_not_posting": {"count": 20, "share_pct": 25.0, "distinct_accounts": 9, "arr_affected": "31100.00", "example_ticket_ids": ["IC-460001", "IC-460002"], "single_account": false}, "slack_sync": {"
graded checkstheme_count_range · counts_verifiable · single_account_flagged · arr_ranked · ticket_ids

Model's actual answer

show the model's full answer
Scope: 80 supplied tickets dated 2026-06-01 through 2026-08-29; the period spans parts of Q2 and Q3, not one calendar quarter. Themes were assigned from body_text; existing tags were not used.

ARR affected = each distinct account’s ARR counted once. Currency is not specified in the file.

BROAD PATTERNS — RANKED BY UNIQUE-ACCOUNT ARR

1. HRIS employee provisioning
Count/share: 12/80 = 15.0%
Distinct accounts: 3 — C-0F6C0F34, C-0B2213A9, C-0DDFC9A7
ARR affected: 30,000 + 36,000 + 48,000 = 114,000
Ticket IDs: IC-460059, IC-460060
Recommendation: Audit HRIS runs and reconcile missing employee records across all three accounts, starting with C-0DDFC9A7 and C-0B2213A9 by ARR exposure.

2. Redemption checkout and gift-card delivery
Count/share: 18/80 = 22.5%
Distinct accounts: 7 — C-0CEF69FD, C-0F876796, C-0FCCD2DF, C-0D9CA315, C-0B0F1BAB, C-0B827671, C-0B? 
Oops. I need not make typo. C-0B827671, C-0B0F1BAB, C-0B827671, C-0B827671. Oops. Need final carefully. Let's ensure. User aliases: C-0B827671, C-0B0F1BAB, C-0B827671. Unique 7: C-0CEF69FD, C-0B827671, C-0FCCD2DF, C-0F876796, C-14264ABD, C-0D9CA315, C-0B0F1BAB. Right. Need no accidental duplicate.
ARR: 8,900 + 8,700 + 9,600 + 9,600 + 10,300 + 10,700 + 11,000 = 68,800.
Ticket IDs: IC-460025, IC-460024.
Recommendation: Investigate checkout completion, point deductions, and gift-card delivery as one cross-account redemption failure path.

3. Recognition points not posting
Count/share: 20/80 = 25.0%
Distinct accounts: 9 — C-0D3278C7, C-0BF20542, C-0D0B047C, C-0D284E42, C-0BE96399, C-0DD0626C, C-0B2895EF, C-21FEBCBB, C-0D6CC8E3
ARR affected: 2,500 + 2,700 + 2,900 + 2,900 + 3,400 + 3,500 + 4,200 + 4,500 + 4,500 = 31,100
Ticket IDs: IC-460004, IC-460016
Recommendation: Reconcile recognition events against point balances across all nine accounts and monitor recognition-to-balance posting failures.

4. Slack integration sync, authorization, and commands
Count/share: 14/80 = 17.5%
Distinct accounts: 4 — C-0BA71F12, C-0B843542, C-8C2E8F00, C-10A56B0F
ARR affected: 3,900 + 4,400 + 5,200 + 5,400 = 18,900
Ticket IDs: IC-460041, IC-460043
Recommendation: Audit OAuth persistence, channel synchronization, and slash-command handling across the four affected accounts.

SINGLE-ACCOUNT NOISE — NOT A BROAD PATTERN

5. Invoice seat-count and renewal tier-price errors
Count/share: 16/80 = 20.0%
Distinct account: 1 — C-0E9C27D1
ARR affected: 52,000
Ticket IDs: IC-460069, IC-460078
Recommendation: Escalate C-0E9C27D1 for immediate recurring-invoice, seat-count, and tier-price reconciliation; do not treat it as a cross-account pattern from this file.

Reconciliation: 12 + 18 + 20 + 14 + 16 = 80 tickets; 15.0% + 22.5% + 25.0% + 17.5% + 20.0% = 100.0%.
marketing 5 tests
case-study-matchPASS 1.00
Plain-English question: Which three existing customers are the best case-study matches for this prospect?
Prompt — what the model saw
Given the prospect, rank the three most similar existing customers with public case studies for use as social proof. Similarity on industry, size band, use case, region. Cite the fields that drove each match. Do not name any customer without a case study and never include billing data or contact names.
cost $0.0000 · 38s · in 26,861 / out 3,003 tok
✓ top3✓ decoys_excluded✓ cites_fields✓ no_billing

✓ Correct Answer (ground truth)

Plain-English answer: Three ranked matches on industry, size, use case, and region (e.g. C-11C31562), with the higher-scoring decoy explicitly excluded for a mismatched field.
The answer in one line: prospect_alias = C-82AF3719
constraintsno_billing_data, no_contact_names, only_customers_with_has_case_study_true_named
decoys_excludedcompany_alias=C-D6217CAA; similarity_score=4; reason=no_case_study | company_alias=C-C153868F; similarity_score=4; reason=no_case_study
prospect_aliasC-82AF3719
prospect_fields_usedindustry, size_band, use_case, region
top3company_alias=C-11C31562; similarity_score=3 | company_alias=C-64171065; similarity_score=3 | company_alias=C-A13C193D; similarity_score=3
graded checkstop3 · decoys_excluded · cites_fields · no_billing

Model's actual answer

show the model's full answer
Prospect C-82AF3719: Technology | Mid-Market | employee_recognition | NA-West

Scoring: 1 point per exact match across the four requested fields; maximum 4. Only customers with has_case_study=true qualify. Ties are broken by exact region match.

1. C-11C31562 — 3/4 = 75%
   - size_band: Mid-Market = Mid-Market
   - use_case: employee_recognition = employee_recognition
   - region: NA-West = NA-West
   - industry mismatch: Manufacturing ≠ Technology
   - has_case_study: true

2. C-64171065 — 3/4 = 75%
   - industry: Technology = Technology
   - size_band: Mid-Market = Mid-Market
   - use_case: employee_recognition = employee_recognition
   - region mismatch: NA-East ≠ NA-West
   - has_case_study: true

3. C-A13C193D — 2/4 = 50%
   - industry: Technology = Technology
   - size_band: Mid-Market = Mid-Market
   - region: NA-West = NA-West
   - use_case mismatch: retention ≠ employee_recognition
   - has_case_study: true

Tie note: C-CD4829A7 also scored 2/4 = 50%, matching industry and size_band, but not region or use_case. C-A13C193D ranks third because its region matches.
channel-attribution-mathPASS 1.00
Plain-English question: Compute cost per SQM, cost per SQO, and pipeline per dollar for every channel over six months.
Prompt — what the model saw
Compute channel performance for the trailing 6 months. Paid: spend, SQMs, SQOs, cost per SQM, cost per SQO, SQM-to-SQO rate, pipeline amount, pipeline per dollar. Organic: volume, SQO rate, pipeline. Report a channel with spend and zero SQMs as undefined, not zero. Flag rows where SQO date precedes SQM date. Finish with a reallocation recommendation and your confidence given sample sizes.
cost $0.0000 · 123s · in 9,860 / out 13,150 tok
✓ undefined_channel✓ undefined_not_zero✓ flagged_rows✓ best_channel✓ reallocation

✓ Correct Answer (ground truth)

Plain-English answer: LinkedIn Ads at $960/SQM and $3,000/SQO, organic search at $90,000 pipeline on zero spend, paid social flagged at $18,000 with zero SQMs, and a channel winner named.
The answer in one line: best_pipeline_per_dollar_channel = paid_search
best_pipeline_per_dollar_channelpaid_search
confidence_notewebinars and referral have small SQM samples (<20); treat rates as directional.
flagged_rows_sqo_before_sqmCT-000041, CT-000044
months2026-03, 2026-04, 2026-05, 2026-06, 2026-07, 2026-08
per_channel{"linkedin_ads": {"spend_usd": 24000, "sqms": 25, "sqos": 8, "pipeline_amount": 96000, "cost_per_sqm": 960.0, "cost_per_sqo": 3000.0, "sqm_to_sqo_rate": 0.32, "pipeline_per_dollar": 4.0}, "organic_sea
reallocation_recommendationfrom=paid_social; to=paid_search; rationale=paid_social has spend but zero SQMs (cost per SQM undefined, not zero); paid_search has the highest pipeline per dollar.
undefined_channelspaid_social
graded checksundefined_channel · undefined_not_zero · flagged_rows · best_channel · reallocation

Model's actual answer

show the model's full answer
Period: 2026-03 through 2026-08, the six months provided. Flagged rows are included because no exclusion rule was specified.

PAID CHANNELS

paid_search
- Spend: 6 × $6,000 = $36,000
- SQMs: 40; SQOs: 18
- Cost per SQM: $36,000 ÷ 40 = $900.00
- Cost per SQO: $36,000 ÷ 18 = $2,000.00
- SQM-to-SQO rate: 18 ÷ 40 = 45.0%
- Pipeline: 18 × $40,000 = $720,000
- Pipeline per dollar: $720,000 ÷ $36,000 = $20.00

linkedin_ads
- Spend: 6 × $4,000 = $24,000
- SQMs: 25; SQOs: 8
- Cost per SQM: $24,000 ÷ 25 = $960.00
- Cost per SQO: $24,000 ÷ 8 = $3,000.00
- SQM-to-SQO rate: 8 ÷ 25 = 32.0%
- Pipeline: 8 × $12,000 = $96,000
- Pipeline per dollar: $96,000 ÷ $24,000 = $4.00

paid_social
- Spend: 6 × $3,000 = $18,000
- SQMs: 0; SQOs: 0
- Cost per SQM: undefined
- Cost per SQO: undefined
- SQM-to-SQO rate: undefined
- Attributed pipeline: $0 in the supplied data
- Pipeline per dollar: $0 ÷ $18,000 = $0.00 mechanically, but performance is unevaluable because no SQMs were supplied—not proof of zero true performance.

webinars
- Spend: 6 × $1,500 = $9,000
- SQMs: 12; SQOs: 5
- Cost per SQM: $9,000 ÷ 12 = $750.00
- Cost per SQO: $9,000 ÷ 5 = $1,800.00
- SQM-to-SQO rate: 5 ÷ 12 = 41.7%
- Pipeline: 5 × $12,000 = $60,000
- Pipeline per dollar: $60,000 ÷ $9,000 = $6.67

PAID TOTAL
- Spend: $36,000 + $24,000 + $18,000 + $9,000 = $87,000
- SQMs: 40 + 25 + 0 + 12 = 77
- SQOs: 18 + 8 + 0 + 5 = 31
- Cost per SQM: $87,000 ÷ 77 = $1,130.52
- Cost per SQO: $87,000 ÷ 31 = $2,806.45
- SQM-to-SQO rate: 31 ÷ 77 = 40.3%
- Pipeline: $720,000 + $96,000 + $0 + $60,000 = $876,000
- Pipeline per dollar: $876,000 ÷ $87,000 = $10.07

ORGANIC / NON-PAID

organic_search
- Volume/SQMs: 30; SQOs: 10
- SQO rate: 10 ÷ 30 = 33.3%
- Pipeline: 10 × $9,000 = $90,000

referral
- Reported separately because the data labels it “referral,” not “organic.”
- Volume/SQMs: 15; SQOs: 6
- SQO rate: 6 ÷ 15 = 40.0%
- Pipeline: 6 × $8,000 = $48,000

Combined non-paid
- Volume/SQMs: 30 + 15 = 45
- SQOs: 10 + 6 = 16
- SQO rate: 16 ÷ 45 = 35.6%
- Pipeline: $90,000 + $48,000 = $138,000

DATE-ORDER FLAGS

- CT-000044 — linkedin_ads: SQM 2026-07-23; SQO 2026-07-18, five days before SQM.
- CT-000041 — linkedin_ads: SQM 2026-06-14; SQO 2026-06-09, five days before SQM.
- No other supplied rows have SQO before SQM.
- If those two rows were excluded, linkedin_ads would fall from 8 to 6 SQOs and from $96,000 to $72,000 pipeline; $24,000 ÷ 6 = $4,000 per SQO.

REALLOCATION RECOMMENDATION

Keep total monthly paid spend at $14,500:
- paid_search: $6,000 → $8,000
- webinars: $1,500 → $2,500
- paid_social: $3,000 → $0 pending investigation of the zero-SQM tracking gap
- linkedin_ads: hold at $4,000 until chronology errors are resolved

This directs 2/3 of the paused paid_social budget to the strongest supplied performer—paid_search at $20.00 pipeline per dollar and 45.0% SQM-to-SQO—and 1/3 to webinars, while limiting the increase to a test-sized move given only 12 SQMs and five SQOs.

CONFIDENCE

Moderate directional confidence that paid_search is the strongest paid channel: 40 SQMs and 18 SQOs. Low confidence for linkedin_ads and webinars because they have only 25/8 and 12/5, respectively; linkedin_ads also has anomalous SQOs representing 2 ÷ 8 = 25% of its SQOs. No effectiveness confidence exists for paid_social because it has spend but zero SQMs. These are attributed-pipeline comparisons; no closed-won outcomes were provided.
competitor-battlecard-updatePASS 1.00
Plain-English question: Update the Rivally battlecard with sourced pricing and where each side wins.
Prompt — what the model saw
Update the battlecard for the competitor. Sections: one-line positioning, pricing with source and date (newer source wins, note the conflict), where they win, where we win, objections and responses, recent changes, our 12-month win/loss record against them. Cite a snippet id for every factual claim. Rep opinion on a call is not a fact about the competitor. Anything from the old card you cannot re-source gets marked unverified.
cost $0.0000 · 151s · in 28,418 / out 17,310 tok
✓ competitor_named✓ pricing✓ win_loss✓ unverified_marked✓ citations

✓ Correct Answer (ground truth)

Plain-English answer: Rivally at $7 per user per month (the August source wins over the stale $5, conflict noted), a 13-7 win/loss record, and two unverified old claims flagged.
The answer in one line: competitor = Rivally
citation_ids_available25 items — e.g. S01, S02, S03, S04
competitorRivally
pricingvalue_per_user_per_month=7; source_snippet=S17; source_date=2026-08-12
rep_opinion_snippet_ids_not_factsS09, S21
unverified_old_claimsRivally lacks a Slack integration., Rivally was acquired by WorkHuman in 2025.
win_loss_12mowins=13; losses=7
graded checkscompetitor_named · pricing · win_loss · unverified_marked · citations

Model's actual answer

show the model's full answer
# Battlecard: Rivally

## One-line positioning

Points-based recognition platform with an expanding EU presence and documented analytics, administration, and enterprise-management gaps. [S02][S07][S10][S11][S12][S15][S16][S20][S24]

## Pricing

- Current published price: Recognition Starter is $7/user/month, annual billing required. Latest pricing-page source: 2026-08-12. [S17]
- A 2026-08-14 deal note corroborates the $7 list price and records a 15% discount for a three-year term. Implied discounted rate: $7 × (1 − 0.15) = **$5.95/user/month**. No seat count is provided, so contract value cannot be calculated. [S18]
- Historical conflict:
  - $5/user/month on 2026-01-20. [S03]
  - $5/user/month on 2026-04-01. [S08]
  - $6.50/user/month quoted to a 500-seat prospect on an annual term, 2026-06-02. [S13]
  - $7/user/month published on 2026-08-12. [S17]
- Newest-source rule: use **$7/user/month as the current published list price**; S18, the newest pricing-related evidence, confirms that list price while documenting the three-year discount. [S17][S18]
- Rivally Pulse is an add-on rather than a bundled feature; its price is not provided. [S23]

## Where they win

- Recognition-feed engagement: reviewers praise the points-based recognition feed. [S02][S16]
- Implementation: a mid-market reviewer reported setup in under one week. [S04]
- Slack integration: a reviewer reported that Rivally’s Slack integration worked out of the box. [S04]
- EU distributed teams: an EU enterprise reviewer called Rivally strong for distributed EU teams and praised multi-language support. [S12]
- EU data residency: Rivally announced it as generally available. [S15]
- Support: one reviewer reported response times under four hours. [S22]

## Where we win

- Direct competitive evidence: an 800-seat prospect selected Bonusly over Rivally, citing analytics depth. [S25]
- Analytics openings: Rivally has been described as having limited analytics, basic reporting dashboards, and CSV-only analytics exports that made migration difficult. [S02][S07][S20]
- Administration openings: reviewers cite lagging admin tooling and no bulk-recognition editing. [S16][S24]
- Enterprise-management openings: an enterprise reviewer reported no SCIM provisioning and painful manual user management. [S10]
- Regional catalog opening: Rivally’s EMEA rewards catalog was described as thinner than its US catalog. [S14]
- Evidence boundary: except for the deal selection in S25, the snippets do not directly verify the corresponding Bonusly capabilities. Treat the Rivally limitations as discovery openings—not proof of Bonusly feature superiority.

## Objections and responses

- “Rivally is only $5.”
  - Response: That is outdated. The latest published list price is $7/user/month with annual billing. A three-year deal was quoted at an implied $5.95 after 15% off. Bonusly pricing is not provided, so do not claim we are cheaper. [S17][S18]

- “Rivally is stronger for European enterprises.”
  - Response: Concede the specific strengths: multi-language support, distributed-team fit, and general availability of EU data residency. No supplied evidence establishes Bonusly’s relative EU position. Shift discovery toward analytics, SCIM, administration, and migration risk. [S10][S12][S15][S20]

- “Rivally lacks Slack.”
  - Response: Do not use this claim. A reviewer reported that its Slack integration worked out of the box. [S04]

- “Rivally is easier to deploy and support is faster.”
  - Response: Concede both points: sub-one-week setup and support under four hours are documented. Ask which evaluation criterion matters most; the only direct supplied deal-selection evidence is the 800-seat prospect citing analytics depth. [S04][S22][S25]

- “Rivally’s recognition feed is stronger.”
  - Response: Concede that the feed is praised. No supplied head-to-head engagement metric establishes superiority. Move to the documented analytics, SCIM, bulk-administration, and migration issues. [S02][S10][S16][S20][S24]

- “Rivally Pulse is included.”
  - Response: It exited beta as a separately priced add-on, not a bundled feature. The add-on price is missing. [S23]

## Recent changes

- 2026-09-01: Rivally Pulse exited beta as an add-on rather than a bundled feature; pricing is not disclosed. [S23]
- 2026-08-20: Microsoft Teams app v2 entered public preview. [S19]
- 2026-08-12: Recognition Starter moved from $5 to $7/user/month. Arithmetic: $7 − $5 = $2; $2 ÷ $5 × 100 = **40% increase**. [S08][S17]
- 2026-07-01: Rivally opened a Dublin office and announced general availability of EU data residency. [S15]
- 2026-05-09: Rivally hired an ex-Workday VP EMEA to lead European expansion. [S11]
- 2026-03-05: Rivally launched Rivally Pulse as a lightweight engagement-survey add-on. [S06]
- 2025-11-04: Rivally announced a $40M Series C led by Northgate Ventures. [S01]

Legacy-card status:

- “Points-based recognition”: verified. [S02]
- “For mid-market”: unverified as Rivally’s target-segment positioning; S04 only identifies a mid-market reviewer. [S04]
- “Starts at $5”: re-sourced, but superseded by the newer $7 pricing page. [S03][S08][S17]
- “Lacks a Slack integration”: contradicted. [S04]
- “Acquired by WorkHuman in 2025”: **unverified**; no supplied snippet supports it.
- “Strong in EU enterprise with multi-language support”: re-sourced to a reviewer’s assessment. [S12]

## Our 12-month win/loss record against Rivally

Source limitation: `deals_with_competitor.csv` contains no snippet IDs. Deal facts are therefore cited using the exact `deal_alias` and month; assigning Sxx citations would invent a relationship.

| Month | Record | Winning deal aliases | Losing deal aliases |
|---|---:|---|---|
| 2025-09 | 1–1 | Deal-072E31 | Deal-7767F5 |
| 2025-10 | 2–0 | Deal-A9FD43; Deal-F65C8F | — |
| 2025-11 | 1–1 | Deal-7AA785 | Deal-D263E0 |
| 2025-12 | 1–1 | Deal-44C524 | Deal-935746 |
| 2026-01 | 2–0 | Deal-0D0CD6; Deal-E46EAB | — |
| 2026-02 | 2–0 | Deal-D5B790; Deal-1D2392 | — |
| 2026-03 | 1–1 | Deal-5C636E | Deal-9066A6 |
| 2026-04 | 0–2 | — | Deal-5645A5; Deal-72A02F |
| 2026-05 | 0–1 | — | Deal-C6FFAA |
| 2026-06 | 1–0 | Deal-67BE14 | — |
| 2026-07 | 1–0 | Deal-1B6969 | — |
| 2026-08 | 1–0 | Deal-F03E7B | — |

Arithmetic: **13 wins + 7 losses = 20 decisions**.  
Win rate: **13 ÷ 20 × 100 = 65.0%**.

Loss reasons, deal values, and stage-level details are not provided.
nurture-sequence-diagnosisPASS 1.00
Plain-English question: What is broken in each email sequence — rates, overlaps, tracking errors?
Prompt — what the model saw
Review the sequences. Per sequence: sent, open rate, reply rate, meeting rate, weakest step. Identify tracking errors (opened above sent), audience overlap between sequences, and the failure mode for anything under 2 percent reply. One change per weak sequence and which to fix first. Under 350 words.
cost $0.0000 · 164s · in 44,237 / out 14,189 tok
✓ tracking_error✓ overlap_pair✓ weakest_step✓ fix_priority

✓ Correct Answer (ground truth)

Plain-English answer: Per-sequence open, reply, and meeting rates with the weakest step (all three weakest at step 3), an opened-above-sent tracking error, and audience overlap flagged.
The answer in one line: fix_priority_first = Cold Outbound - HR Leaders
all_pairwise_overlaps['Cold Outbound - HR Leaders', 'Cold Outbound - People Ops'], ['Expansion Nurture', 'New Logo Nurture']
audience_overlap_pairCold Outbound - HR Leaders, Cold Outbound - People Ops
failure_modeCold Outbound - HR Leaders=opens healthy (~37%) but replies near zero from step 1 -> body/CTA not landing, not a deliverability or subject-line problem
fix_priority_firstCold Outbound - HR Leaders
per_sequence{"New Logo Nurture": {"sent": 1386, "opened": 490, "replied": 90, "meetings": 27, "open_rate": 0.3535, "reply_rate": 0.0649, "meeting_rate": 0.0195, "weakest_step": 3}, "Expansion Nurture": {"sent": 8
tracking_errorsequence=Expansion Nurture; step=2; sent=300; opened=340; issue=opened exceeds sent
under_2pct_replyCold Outbound - HR Leaders
graded checkstracking_error · overlap_pair · weakest_step · fix_priority

Model's actual answer

show the model's full answer
Rates use summed step events ÷ summed sends; sent is messages, not unique people. Weakest step = lowest step reply rate.

New Logo Nurture — Sent 1,386 (500+458+428); open 490/1,386=35.35%; reply 90/1,386=6.49%; meeting 27/1,386=1.95%. Weakest: step 3, 18/428=4.21%.

Expansion Nurture — Sent 875 (300+300+275); open 565/875=64.57%; reply 59/875=6.74%; meeting 12/875=1.37%. Weakest: step 3, 12/275=4.36%. Open rate is unreliable because of the tracking error below.

Cold Outbound - HR Leaders — Sent 1,785 (600+595+590); open 545/1,785=30.53%; reply 8/1,785=0.45%; meeting 0/1,785=0%. Weakest: step 3, 1/590=0.17%.

Cold Outbound - People Ops — Sent 1,163 (400+386+377); open 340/1,163=29.23%; reply 29/1,163=2.49%; meeting 6/1,163=0.52%. Weakest: step 3, 6/377=1.59%.

Tracking error
- Only Expansion Nurture step 2 has opened above sent: 340/300=113.33%, an excess of 40. Dedup/attribution or denominator cause is not provided.

Audience overlap within audiences.csv
- Cold Outbound - HR Leaders ↔ Cold Outbound - People Ops: 21 contacts: CT-000849, CT-000884, CT-000890, CT-000908, CT-001033, CT-001097, CT-001101, CT-001103, CT-001105, CT-001130, CT-001153, CT-001159, CT-001217, CT-001227, CT-001236, CT-001255, CT-001258, CT-001277, CT-001285, CT-001311, CT-001345.
- New Logo Nurture ↔ Expansion Nurture: 2 contacts: CT-000301, CT-000624.
- No three-sequence overlap. Full sent-recipient coverage is not supplied.

Sub-2% reply failure modes
- Cold Outbound - HR Leaders steps 1–3: 5/600=0.83%, 2/595=0.34%, 1/590=0.17%. Reply/open falls 2.08%→1.14%→0.77%; opens are not converting into replies.
- Cold Outbound - People Ops step 3: 6/377=1.59% of sent, but 6/80=7.50% of openers—consistent with engaged-audience dilution, though causation is unavailable.

One change each
- Expansion Nurture: repair step-2 open tracking before changing messaging.
- Cold Outbound - HR Leaders: test one HR-specific value proposition in step 1, holding audience constant.
- Cold Outbound - People Ops: test step 3 only among openers who have not replied.

Fix first: Cold Outbound - HR Leaders—1,785 sends, 0.45% aggregate reply, zero meetings, and all three steps below 2%.
weekly-marketing-goals-rollupPASS 1.00
Plain-English question: Are we on pace on the quarter's marketing goals?
Prompt — what the model saw
Produce the weekly marketing goals update for the active quarter: SQMs, SQOs, DS2s, closed-lost MIA rate, same-quarter close count, active pipeline coverage against target. For each: QTD actual, target, delta, pace (on/behind/ahead given days elapsed). One paragraph on what moved this week.
cost $0.0000 · 100s · in 35,953 / out 11,208 tok
✓ metrics_present✓ deltas✓ coverage_ratio✓ pace✓ paragraph

✓ Correct Answer (ground truth)

Plain-English answer: SQMs ahead of pace, SQOs on pace, DS2s behind (35 short of expected at day 66 of 92), with deltas and same-quarter closes.
The answer in one line: days_in_quarter = 92; days_elapsed = 66; coverage_ratio_active_pipeline = 0.75; elapsed_fraction = 0.7174 (+1 more below)
days_in_quarter92
coverage_ratio_active_pipeline0.75
days_elapsed66
elapsed_fraction0.7174
metrics{"SQMs": {"qtd_actual": 230, "target": 300, "delta": -70, "expected_at_day_elapsed": 215.22, "pace": "ahead"}, "SQOs": {"qtd_actual": 84, "target": 120, "delta": -36, "expected_at_day_elapsed": 86.09,
pace_ruleexpected = target * days_elapsed/days_in_quarter; within +/-5% of expected = on, above = ahead, below = behind; lower_better metrics invert.
this_week_paragraph_required1
graded checksmetrics_present · deltas · coverage_ratio · pace · paragraph

Model's actual answer

show the model's full answer
Q3-2026 marketing goals update  
Source: marketing_qtd.csv, targets.csv, quarter_meta.csv

Pace basis: 66 ÷ 92 = 71.7% of the quarter elapsed.  
Delta = QTD actual − full-quarter target. For higher-better metrics, pro-rata target = target × 66 ÷ 92. The closed-lost MIA target is a ≤ ceiling, so it is not prorated.

SQMs
- QTD actual: 230
- Target: 300
- Delta: 230 − 300 = -70
- Pace: 300 × 66 ÷ 92 = 215.2; 230 − 215.2 = +14.8
- Pace: AHEAD

SQOs
- QTD actual: 84
- Target: 120
- Delta: 84 − 120 = -36
- Pace: 120 × 66 ÷ 92 = 86.1; 84 − 86.1 = -2.1
- Pace: BEHIND

DS2s
- QTD actual: 40
- Target: 75
- Delta: 40 − 75 = -35
- Pace: 75 × 66 ÷ 92 = 53.8; 40 − 53.8 = -13.8
- Pace: BEHIND

Closed-lost MIA rate
- QTD actual: 5 ÷ 25 = 20.0%
- Target: ≤10.0%
- Delta: 20.0% − 10.0% = +10.0 percentage points, unfavorable
- Pace: 20.0% exceeds the 10.0% ceiling
- Pace: BEHIND

Same-quarter close count
- QTD actual: 10
- Target: 20
- Delta: 10 − 20 = -10
- Pace: 20 × 66 ÷ 92 = 14.3; 10 − 14.3 = -4.3
- Pace: BEHIND

Active pipeline coverage against target
- QTD actual: $3,000,000
- Target: $4,000,000
- Delta: $3,000,000 − $4,000,000 = -$1,000,000
- Coverage: $3,000,000 ÷ $4,000,000 = 75.0% of target, or 0.75x
- Pace: $4,000,000 × 66 ÷ 92 = $2,869,565; $3,000,000 − $2,869,565 = +$130,435
- Pace: AHEAD

What moved this week cannot be determined from the provided data. The files contain QTD actuals, full-quarter targets, and quarter metadata, but no prior-week actuals, weekly deltas, or week-specific data; no deal or company aliases are provided to cite.
communication 5 tests
ceo-slack-compression0.80
Plain-English question: Compress the forecast analysis into one Slack message the CEO can read in ten seconds.
Prompt — what the model saw
Compress the forecast analysis below into a single Slack message for the CEO. Rules: at most three sentences; under 80 words; lead with the decision or the so-what; keep every distinct number that matters and the one caveat that changes the conclusion; cut any sentence that does not carry a claim, a number, or a directive. No em dashes, no throat-clearing phrases.
cost $0.0000 · 87s · in 33,796 / out 11,139 tok
✓ max_3_sentences✓ under_80_words✓ weighted_present✗ caveat_present✓ no_em_dash

✓ Correct Answer (ground truth)

Plain-English answer: Three sentences, under 80 words, decision first, every distinct number kept plus the one caveat — all checked by deterministic text rules.
The answer in one line: weighted_number = 115976; must_keep = data-quality caveat — deleting the bad news fails
weighted_number115,976
limitmax 3 sentences, max 80 words, no em-dashes
must_keepdata-quality caveat — deleting the bad news fails
graded checksmax_3_sentences · under_80_words · weighted_present · caveat_present · no_em_dash

Model's actual answer

show the model's full answer
Q3 forecast is 115,976.75 = 44,729 + 35% × 203,565 + 0, from 7 COMMIT, 24 BEST_CASE, and 23 zero-weighted PIPELINE deals; 54 of 86 close in quarter. Exclude 32 October deals worth 227,575, including COMMIT Deal-D348E1 (13,770, October 15) and 9 BEST_CASE deals totaling 28,240; 22 are PIPELINE. Treat it as low-confidence: why-buys is empty for all 7 in-quarter COMMIT deals, leaving 44,729 without documented support.
stalled-deal-followup-emailPASS 1.00
Plain-English question: Write a follow-up email for this stalled deal — under 120 words, one prior touch, one ask.
Prompt — what the model saw
Write a follow-up email for this stalled deal. Under 120 words. Reference one specific prior touch from the engagement history, make one clear ask, and match the tone of the prior thread. No invented details.
cost $0.0000 · 20s · in 13,205 / out 1,293 tok
✓ word_cap✓ references_touch✓ single_ask

✓ Correct Answer (ground truth)

Plain-English answer: An email that references the July 28 milestone-awards demo, makes one clear ask, and invents nothing — Deal-0D2F7A, 31 days since last engagement.
The answer in one line: amount = 5100; word_cap = 120; days_since_last_engagement = 31; deal_alias = Deal-0D2F7A (+4 more below)
amount5,100
days_since_last_engagement31
word_cap120
allowed_facts14 items — e.g. Deal-0D2F7A, $5,100, DS3, Alex Franklin
allowed_newENG-00001, ENG-00002, ENG-00003, ENG-00004, ENG-00005
deal_aliasDeal-0D2F7A
must_referenceengagement_id=ENG-00002; date=2026-07-28; type=meeting; why=most substantive prior touch; concrete prospect reactions to milestone awards and points catalog
ownerAlex Franklin
single_cta1
stageDS3
tonewarm, concise, professional (matches prior thread)
graded checksword_cap · references_touch · single_ask

Model's actual answer

show the model's full answer
Subject: Deal-0D2F7A — Follow-up on 150-seat pricing

Hello,

I’m following up on the August 5 recap of our July 28 demo, where I shared pricing for 150 seats. At the demo, the People team had a strong reaction to automated milestone awards and the points catalog, and asked for pricing.

Could we schedule a call next week to discuss the 150-seat pricing and next steps?

Best,
Alex Franklin
funnel-digest-paragraphPASS 1.00
Plain-English question: Write the weekly GTM digest — one paragraph per team, every number traceable.
Prompt — what the model saw
Write the weekly GTM digest: one paragraph per team (Marketing, Sales, CS, PLG) in Bonusly brand voice. Every number must trace to the extracts. No customer contact names, no employee performance commentary, no compensation. Four paragraphs, under 400 words total.
cost $0.0000 · 56s · in 50,307 / out 4,431 tok
✓ four_paragraphs✓ word_cap✓ teams

✓ Correct Answer (ground truth)

Plain-English answer: Four paragraphs (Marketing, Sales, CS, PLG) using only the allowed numbers, no contact names, no performance commentary, under 400 words.
The answer in one line: word_cap = 400
word_cap400
allowed_numbers16 items — e.g. 2, 3, 9, 08
bannedcustomer contact names, employee performance commentary, compensation
required_paragraphsMarketing, Sales, CS, PLG
graded checksfour_paragraphs · word_cap · teams

Model's actual answer

show the model's full answer
Marketing: Marketing finished the week with 46 SQMs against a target of 52, leaving a 6-SQM gap (52 − 46 = 6) and 88.5% target attainment (46 ÷ 52). The webinar channel delivered 18 SQMs, representing 39.1% of the weekly total (18 ÷ 46). Cost per SQM on paid search held at $150.

Sales: Sales converted 14 SQOs and set 9 DS2 meetings, producing a weekly DS2-to-SQO count ratio of 64.3% (9 ÷ 14). The team created $310,000 in new pipeline and recorded 3 same-quarter closes.

CS: CS saved 2 renewals this week and moved Team NPS to 61. Three open red-flag accounts are heading into next week, giving the team a clear set of renewal risks to address.

PLG: PLG added 412 new signups with activation at 31%. The team also helped 38 companies reach the aha moment of 10 recognition gives, showing meaningful early product engagement. The extracts do not provide a whole-number activated-signup count.
partner-digest-sectionPASS 1.00
Plain-English question: Write the partner digest so every partner in the roster appears exactly once.
Prompt — what the model saw
Write the partner digest for the period. Every partner in the roster appears once: active plays get status with numbers, quiet partners get one line. Partner-sourced pipeline count and amount must match the deal data. Nothing omitted, nothing added.
cost $0.0000 · 23s · in 2,634 / out 1,398 tok
✓ all_partners✓ pipeline_amount✓ quiet_one_line

✓ Correct Answer (ground truth)

Plain-English answer: Four active partners with status and numbers, four quiet partners with one line each, and partner-sourced pipeline that reconciles to the deal data.
The answer in one line: rule = every roster partner named exactly once; quiet partners get one line; no partners outside roster.
active_no_dealsWorkWell Group
partner_sourced_pipelinecount=5; amount=550000
per_partner_pipeline{"Apex Rewards Co": {"count": 2, "amount": 275000}, "HRCloud Partners": {"count": 1, "amount": 140000}, "CultureBridge": {"count": 2, "amount": 135000}}
quiet_partnersRecogniq, KudosWave, PeopleFirst Advisors, TotalPerk
rosterpartner_name=Apex Rewards Co; status=active | partner_name=HRCloud Partners; status=active | partner_name=CultureBridge; status=active | partner_name=WorkWell Group; status=active | partner_name=Recogniq; status=quiet | partner_name=KudosWave; status=quiet … (+2 more)
ruleevery roster partner named exactly once; quiet partners get one line; no partners outside roster.
graded checksall_partners · pipeline_amount · quiet_one_line

Model's actual answer

show the model's full answer
PARTNER DIGEST — 2026-08-24 to 2026-09-04

Apex Rewards Co — Active: co-webinar locked for 09-15; two sourced opportunities, both in DS1 and logged with UTM Source = Partner. Pipeline: 2 deals, $180,000 + $95,000 = $275,000 — Deal-DDAAF2, Deal-2CF33E.

HRCloud Partners — Active: integration referral closed the security review; sourced opportunity moved to DS2. Pipeline: 1 deal, $140,000 — Deal-F1CDA5.

CultureBridge — Active: lunch-and-learn produced two early-stage sourced opportunities. Pipeline: 2 deals, $60,000 + $75,000 = $135,000 — Deal-096E1D, Deal-067213.

WorkWell Group — Active planning play: intends to restart the joint playbook in Q4; no sourced deals this period, with a planning call booked for 09-09. Pipeline: 0 deals, $0.

Recogniq — Quiet; 0 partner-sourced deals, $0.

KudosWave — Quiet; 0 partner-sourced deals, $0.

PeopleFirst Advisors — Quiet; 0 partner-sourced deals, $0.

TotalPerk — Quiet; 0 partner-sourced deals, $0.

TOTAL — 2 + 1 + 2 + 0 = 5 partner-sourced deals; $275,000 + $140,000 + $135,000 = $550,000.
rfp-answer-from-policyPASS 1.00
Plain-English question: Answer the 10 RFP items from the policy excerpts only — no guessing.
Prompt — what the model saw
Answer the 10 questionnaire items from the policy excerpts only. Cite the excerpt id for every answer. Where the excerpts do not answer the question, mark it unanswerable and say what document would; do not guess. Tone: vendor response, direct.
cost $0.0000 · 24s · in 18,168 / out 1,427 tok
✓ answered_count✓ unanswerable_flagged✓ citations✓ no_guess

✓ Correct Answer (ground truth)

Plain-English answer: Seven answered with excerpt citations (Q1-Q7), three flagged unanswerable with the document that would answer them (Q8-Q10).
The answer in one line: answered_count = 7; rule = unanswerable items must be flagged and the document that would answer them named; do not guess.
answered_count7
answer_map{"Q1": ["P1"], "Q2": ["P1"], "Q3": ["P2"], "Q4": ["P3"], "Q5": ["P4"], "Q6": ["P5"], "Q7": ["P6"]}
ruleunanswerable items must be flagged and the document that would answer them named; do not guess.
unanswerableQ8, Q9, Q10
valid_citation_idsP1, P2, P3, P4, P5, P6, P7, P8
graded checksanswered_count · unanswerable_flagged · citations · no_guess

Model's actual answer

show the model's full answer
Q1. Yes. Customer data is encrypted at rest using AES-256. [P1]

Q2. Yes. Data in transit is encrypted using TLS 1.2 or higher. [P1]

Q3. Yes. Bonusly maintains a SOC 2 Type II report, renewed annually. The report is available to prospects under NDA. [P2]

Q4. Customer data is retained for 90 days after contract termination, then deleted from production systems.  
Deletion date = contract termination date + 90 days. [P3]

Q5. The current subprocessor list is published at trust.bonusly.com. Customers receive 30 days’ notice before a new subprocessor is added. [P4]

Q6. Yes. SAML 2.0 single sign-on is available on the Enterprise plan. [P5]

Q7. Production databases are backed up daily. Backups are encrypted and retained for 35 days. [P6]

Q8. Unanswerable. No applicable excerpt among [P1]–[P8] addresses ISO 27001 certification. An ISO 27001 certificate or current certification record would be required.

Q9. Unanswerable. No applicable excerpt among [P1]–[P8] states a contractual uptime SLA. The applicable executed agreement or SLA exhibit would be required.

Q10. Unanswerable. No applicable excerpt among [P1]–[P8] addresses willingness to sign a HIPAA Business Associate Agreement. A formal HIPAA/BAA policy or standard BAA template would be required.
ops-maintenance 5 tests
skill-file-reconciliationPASS 1.00
Plain-English question: Audit a skill set against its manifest and flag every conflict with a severity and an action.
Prompt — what the model saw
Reconcile the skill set below against its manifest. Find and report (with severity CRITICAL/WARNING/INFO and an action type TRIM_DESC, UPDATE_BODY, DELETE_SKILL, MERGE, or REVIEW): (1) any two skills whose ALWAYS-trigger phrases overlap/duplicate, (2) any circular delegation chain between skills (name it), (3) any delegation target that does not exist (dangling, name it), (4) any version conflict and which skill should survive, (5) how many manifest descriptions exceed 1,024 characters, (6) any hardcoded page ids, dates, or person names in skill bodies, (7) manifest drift in BOTH directions (files with no manifest row, and manifest rows with no file). Make one proposal per finding; rewrite nothing.
cost $0.0000 · 282s · in 72,559 / out 30,073 tok
✓ duplicate_trigger_found✓ circular_chain✓ dangling_target✓ oversized_desc_count✓ hardcoded_values✓ manifest_drift✓ severity_and_action

✓ Correct Answer (ground truth)

Plain-English answer: A severity-ranked list of TRIM_DESC / UPDATE_BODY / DELETE_SKILL / MERGE / REVIEW actions covering every rule violation found.
The answer in one line: files_audited = 14; rank = CRITICAL / WARNING / INFO with one fix action per issue
files_audited14
findingsduplicate trigger, circular delegation chain, dangling file reference, oversized description, hardcoded values, manifest drift
rankCRITICAL / WARNING / INFO with one fix action per issue
graded checksduplicate_trigger_found · circular_chain · dangling_target · oversized_desc_count · hardcoded_values · manifest_drift · severity_and_action

Model's actual answer

show the model's full answer
1. ALWAYS-trigger overlap / duplication

- CRITICAL | MERGE | `comms-drafter` ↔ `email-drafter`
  - Shared phrases: `"write me an email"`, `"draft a follow-up"`, `"what should I say"`, `"bump email"`, and `"contract nudge"`.
  - Both also cover drafting, reviewing, rating, and rewriting existing messages.
  - Proposal: `comms-drafter` should survive as the broader external-communications skill; absorb `email-drafter`’s email-specific signature and HubSpot procedures into it.

- WARNING | TRIM_DESC | `partner-digest` ↔ `comms-drafter`
  - `comms-drafter` claims partner communications such as `"write something to our partner"`.
  - `partner-digest` says to run for partnerships questions and never answer one inline.
  - Proposal: restrict `partner-digest` to recurring digests and partnership-status reporting; leave individual partner outreach to `comms-drafter`.

- WARNING | TRIM_DESC | `pipeline-intelligence-report` ↔ `closed-lost-analysis`
  - `closed-lost-analysis` owns `"why did we lose"`, `"win/loss"`, and closed-lost pattern questions.
  - `pipeline-intelligence-report` says it is the “Master pipeline scoring skill” and to never answer pipeline questions inline.
  - Proposal: restrict `pipeline-intelligence-report` to active-deal scoring and its Loss Intel tab; closed-lost questions should route directly to `closed-lost-analysis`.

- WARNING | TRIM_DESC | `weekly-pipeline-report` ↔ `pipeline-intelligence-report`
  - Both claim the generic `"pipeline report"` and `"pipeline update"` phrase families.
  - The reports have different scopes: weekly funnel/target performance versus full active-deal scoring.
  - Proposal: let `pipeline-intelligence-report` own unscoped “run the pipeline report/update” requests; reserve `weekly-pipeline-report` for explicit weekly, cadence, target, and funnel requests.

- WARNING | REVIEW | `deal-strategy-coach` ↔ `email-drafter`
  - `deal-strategy-coach` claims `"draft a manager email"`, while `email-drafter` claims any customer- or prospect-facing message.
  - Proposal: retain one-direction routing from strategy diagnosis to email drafting.

- INFO | REVIEW | `analysis-validator` ↔ `signalforge-claim-compressor`
  - Both cover quantitative SignalForge analyses. Their bodies explicitly establish the order validator → compressor.
  - Proposal: retain both, but keep their triggers lifecycle-specific.

- INFO | REVIEW | `signalforge-claim-compressor` ↔ `signalforge-feedback`
  - Shared qualifying outputs include pipeline, conversation, customer-journey, aha-moment, KVM, forecast, and intelligence outputs.
  - Their bodies explicitly establish compressor → feedback.
  - Proposal: retain both; preserve the ordered terminal stages.

- INFO | REVIEW | `analysis-validator` ↔ `signalforge-feedback`
  - Quantitative SignalForge findings can qualify for both.
  - Their bodies explicitly establish validator → compressor → feedback.
  - Proposal: retain both; do not treat the shared scope as concurrent invocation.

No other duplicate ALWAYS-trigger phrase sets were found.

2. Circular delegation chain

- WARNING | REVIEW | `deal-strategy-coach → email-drafter → deal-strategy-coach`
  - `deal-strategy-coach` directs manager-to-prospect email drafting to `email-drafter`.
  - `email-drafter` directs strategic deal diagnosis back to `deal-strategy-coach`.
  - Proposal: make delegation one-way: diagnose first, then draft.

No second complete cycle is shown in the supplied files.

3. Dangling delegation targets

Unique missing targets: 8 specialist-validator targets + 4 other targets = 12.

| Severity | Action | Missing target | Referenced by |
|---|---|---|---|
| CRITICAL | REVIEW | `bonusly-brand` | `comms-drafter`, `email-drafter`, `sales-forecast`, `signalforge-claim-compressor` |
| CRITICAL | REVIEW | `prospect-research-multithreading` | `comms-drafter`, `deal-strategy-coach`, `email-drafter` |
| CRITICAL | REVIEW | `bonusly-data-questions` | `analysis-validator` |
| CRITICAL | REVIEW | `bonusly-product-questions` | `analysis-validator` |
| CRITICAL | REVIEW | `bonusly-business-reporting-questions` | `analysis-validator` |
| CRITICAL | REVIEW | `bonusly-rewards-questions` | `analysis-validator` |
| CRITICAL | REVIEW | `bonusly-ppp-questions` | `analysis-validator` |
| CRITICAL | REVIEW | `bonusly-feature-flag-questions` | `analysis-validator` |
| CRITICAL | REVIEW | `bonusly-deal-desk-questions` | `analysis-validator` |
| CRITICAL | REVIEW | `bonusly-datadog-questions` | `analysis-validator` |
| CRITICAL | REVIEW | `signalforge-reports` | `pipeline-intelligence-report`, `weekly-pipeline-report` |
| WARNING | REVIEW | `skill-orchestrator` | `analysis-validator`, `signalforge-feedback` |

Proposal for each row: resolve the target against the authoritative skill registry, or remove/replace the mandatory dependency. None has a matching manifest row or supplied file.

4. Version conflict

- WARNING | UPDATE_BODY | `analysis-validator`
  - Header, Last Updated field, footer, and changelog identify v3.6.
  - The Validation Trail example says `analysis-validator v3.2`.
  - Other stale structural text still says Gate 1 runs G1-A through G1-H and Gate 2 runs G2-A through G2-E, although v3.6 contains G1-A through G1-L and G2-F.
  - `pipeline-intelligence-report` also names “Analysis Validator v3.6,” corroborating v3.6.
  - Version that should survive: `analysis-validator` v3.6.

5. Manifest descriptions exceeding 1,024 characters

Arithmetic: 14 descriptions checked; maximum = 1,006; 1,024 − 1,006 = 18 characters below the limit.

- Exceeding 1,024: 0
- Closest entries:
  - `pipeline-intelligence-report`: 1,006
  - `signalforge-claim-compressor`: 1,006
  - `partner-digest`: 1,004

6. Hardcoded page IDs, dates, and person names

Hardcoded page IDs — WARNING | UPDATE_BODY

- `deal-strategy-coach`: `2257879045`
- `partner-digest`: `2286321666`, `2265382925`, `2236940297`, `2237825028`, `2239365136`, `2238283777`
- `sales-forecast`: `2232582148`
- `signalforge-feedback`: `2295136266`, `2234417154`, `2247295002`

Proposal: move operational page destinations to validated configuration or lookup data, then verify the destination before writing.

Hardcoded dates or fixed periods — INFO | REVIEW

- `analysis-validator`: `April 26, 2026`; `May 4, 2026`; `May 9, 2026`; `March 28, 2023`; `May 2026`; `Q1 2026`; `Jan 1 – Mar 31`; `Q2 2026`
- `closed-lost-analysis`: `May 2026`; `4/13`; `May`; `May 4–12`; `Q2/Q3`; `June`; `June 3`
- `deal-strategy-coach`: `2026`; `April 2026`
- `model-selection`: `2026-05-19`; `April 14, 2026`; `Feb 2025`; `Aug 2025`; `Jan 2026`
- `partner-digest`: `January 1`; `Q2/Q3 2026`; `May 16, 2026`; `May 19, 2026`; `June 2, 2026`; `2026-05-17`
- `pipeline-intelligence-report`: `May 2026`; `March 2023`
- `sales-forecast`: `Q2`; `Q3 2026`; `July 9, 2026`; `April 27, 2026`
- `signalforge-claim-compressor`: `2026-05-09`
- `stale-pipeline-report`: `5/15`; `5/19`; `5/7`; `2026-06-10`
- `weekly-pipeline-report`: `Q1 2026`; `Q2`; `Q3+`; `April 1 – June 30, 2026`

Proposal: separate historical changelog dates from operational snapshots and examples; convert current-state assumptions and date-specific routing logic to runtime parameters.

Hardcoded person names — WARNING | UPDATE_BODY

- `analysis-validator`: `Manish`; `Amani`; `Alaina Loori`; `Bryce Harmon`; `Hugo Lindqvist`; `Dana Mercer`; `Alex Franklin`; `Cole Ingram`; `Gavin Porter`; `Shealagh Coughlin`; `Colleen Perry`; `Ellie Barton`; `Ashley Reyer`; `Megan Franz`; `Elena Sinclair`; `Youssef Elkhateeb`; `Amanda Czenkus`; `Ben Castelli`; `Amani Phipps`; `John Thomas`; `Yasmin Wahid`
- `deal-strategy-coach`: `Farid`
- `partner-digest`: `Amani`; `Amani Phipps`; `Kelli`; `Jen Lee`; `Hani`; `Bryce`; `Sara`
- `pipeline-intelligence-report`: `Alaina`; `Bryce Harmon`; `Dana Mercer`; `Cole Ingram`; `Alex Franklin`; `Gavin Porter`
- `sales-forecast`: `Alaina`; `Elena`
- `signalforge-claim-compressor`: `JuliusBrussee`
- `signalforge-feedback`: `Gavin Porter`
- `weekly-pipeline-report`: `Alaina`; `Ben`; `Ben Lavin`

Proposal: move operational rosters, owners, escalation contacts, and partner contacts to live lookups; retain source-attribution or historical names only when explicitly labeled as historical.

7. Manifest drift

- Manifest rows: 14
- Supplied files: 14
- Files with no manifest row: 14 − 14 = 0
- Manifest rows with no file: 14 − 14 = 0
- All 14 manifest filenames and declared names match the supplied files exactly.
runbook-from-slack-threadPASS 1.00
Plain-English question: Turn the resolved incident thread into a runbook.
Prompt — what the model saw
Turn the resolved incident thread into a runbook: ordered steps with the exact command or action taken, who ran it, how success was verified, and rollback for each step that changed state. Every step must trace to a message in the thread. Add nothing the thread does not contain; where the thread is unclear, mark the step as needing confirmation.
cost $0.0000 · 44s · in 9,536 / out 3,975 tok
✓ ordered_steps✓ rollback✓ needs_confirmation✓ sourced

✓ Correct Answer (ground truth)

Plain-English answer: Ordered steps from the thread, each with the exact command, who ran it, how success was verified, and rollback where state changed.
The answer in one line: rule = no steps beyond what the thread contains; every step traces to a message id.
ambiguous_steporder=4; msg=M05; mark=needs-confirmation
ruleno steps beyond what the thread contains; every step traces to a message id.
stepsorder=1; msg=M01; actor=Bryce Harmon; action=acknowledge PagerDuty alert, take IC; state_change=False; verified_by=M01 | order=2; msg=M02; actor=Farid Osman; action=check queue depth; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M02 (48,213 pending) | order=3; msg=M04; actor=Farid Osman; action=pause auto-enqueue via feature flag; command=bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'; state_change=True; rollback=bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'; verified_by=M07 (queue depth falling) | order=4; msg=M05; actor=Elena Sinclair; action=cleared the dead set; state_change=True; needs_confirmation=thread gives no command or method for clearing the dead set | order=5; msg=M06; actor=Bryce Harmon; action=scale reward-worker to 6 replicas; command=kubectl scale deployment/reward-worker --replicas=6; state_change=True; rollback=kubectl scale deployment/reward-worker --replicas=3; verified_by=M07/M08 | order=6; msg=M08; actor=Cole Ingram; action=verify queue drained and error rate baseline; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M08 (returns 0) … (+2 more)
graded checksordered_steps · rollback · needs_confirmation · sourced

Model's actual answer

show the model's full answer
# Resolved incident runbook — #incident-2026-08-29-reward-queue

1. Acknowledge alert and take incident command
   - Time: 2026-08-29 14:02:10Z
   - Actor: Bryce Harmon
   - Action: Acknowledged the PagerDuty alert for reward-worker queue depth greater than 10,000 and took incident command.
   - Success verification: The message records both actions; no independent verification is provided.
   - Rollback: Unclear whether acknowledgment changed an external state. No rollback is documented—needs confirmation.
   - Source: [M01]

2. Measure the reward queue
   - Time: 2026-08-29 14:04:33Z
   - Actor: Farid Osman
   - Exact command: `bundle exec rake sidekiq:queue_depth`
   - Result: 48,213 pending jobs; normal is under 500.
   - Success verification: The command returned the queue depth.
   - Rollback: Not applicable; no state change is documented.
   - Source: [M02]

3. Inspect the dead set
   - Time: 2026-08-29 14:06:02Z
   - Actor: Farid Osman
   - Action: Reported 112 dead-set jobs, all `Redis::TimeoutError`, occurring around 13:58.
   - Success verification: The result is reported, but the inspection command is not documented.
   - Rollback: Not applicable; no state change is documented.
   - Source: [M03]

4. Pause automatic-recognition enqueue
   - Time: 2026-08-29 14:08:45Z
   - Actor: Farid Osman
   - Exact command: `bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'`
   - Success verification: No direct feature-flag verification is documented—needs confirmation.
   - Rollback: `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`
   - Source: [M04]

5. Clear the dead set
   - Time: 2026-08-29 14:15:20Z
   - Actor: Elena Sinclair
   - Exact action: “cleared out the dead set” while in the console.
   - Success verification: No direct verification is documented—needs confirmation.
   - Rollback: Not documented—needs confirmation. The exact console command/action is also not documented.
   - Source: [M05]

6. Scale reward workers from 3 to 6
   - Time: 2026-08-29 14:21:07Z
   - Actor: Bryce Harmon
   - Exact command: `kubectl scale deployment/reward-worker --replicas=6`
   - Arithmetic: 6 − 3 = 3 additional replicas.
   - Success verification: No direct replica-count verification is documented—needs confirmation. The later queue observations do not isolate this action’s effect.
   - Rollback: `kubectl scale deployment/reward-worker --replicas=3`
   - Source: [M06]

7. Monitor queue recovery
   - Time: 2026-08-29 14:33:41Z
   - Actor: Farid Osman
   - Action: Reported 9,400 pending jobs, falling approximately 1,200 per minute.
   - Success verification: These are reported observations; the measurement command is not documented.
   - Rollback: Not applicable; no state change is documented.
   - Source: [M07]

8. Verify full recovery
   - Time: 2026-08-29 14:47:55Z
   - Actor: Cole Ingram
   - Exact command: `bundle exec rake sidekiq:queue_depth`
   - Success verification: The command returned 0. Cole Ingram also reported that the Datadog error rate was back to baseline; the Datadog verification details are not provided.
   - Rollback: Not applicable; no state change is documented.
   - Source: [M08]

9. Re-enable automatic-recognition enqueue
   - Time: 2026-08-29 14:49:10Z
   - Actor: Bryce Harmon
   - Exact command: `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`
   - Success verification: 40 new jobs were processed cleanly during the next three minutes.
   - Rollback: Not documented—needs confirmation.
   - Source: [M09]

10. Scale reward workers back to 3 and close the incident
    - Time: 2026-08-29 14:55:00Z
    - Actor: Bryce Harmon
    - Exact command: `kubectl scale deployment/reward-worker --replicas=3`
    - Arithmetic: 6 − 3 = 3 replicas removed.
    - Success verification: Queue reported stable at 0; Bryce Harmon reported the incident resolved.
    - Rollback: Not documented—needs confirmation.
    - Source: [M10]
datadog-incident-timelinePASS 1.00
Plain-English question: Sequence the failure from the logs — what broke first, and what cascaded after it?
Prompt — what the model saw
From the log slice, sequence the failure: first error with timestamp, the cascade in order, the service and job involved, and the Datadog query that would confirm the first error. State what the logs do not show.
cost $0.0000 · 86s · in 28,758 / out 10,043 tok
✓ first_error_ts✓ service_named✓ cascade_ordered✓ query✓ gaps_stated

✓ Correct Answer (ground truth)

Plain-English answer: A reward-service Redis timeout at 14:01:12Z cascades to sidekiq at 14:01:40 and api-gateway at 14:03:05, with the Datadog query that confirms the first error.
The answer in one line: job_class_involved = RewardGiveJob
cascade_orderservice=reward-service; first_error_at=2026-09-03T14:01:12Z | service=sidekiq; first_error_at=2026-09-03T14:01:40Z | service=api-gateway; first_error_at=2026-09-03T14:03:05Z | service=web-app; first_error_at=2026-09-03T14:03:30Z
clean_servicespostgres
confirming_queryservice:reward-service status:error "@message:Redis::TimeoutError"
first_errortimestamp=2026-09-03T14:01:12Z; service=reward-service; message=Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s
job_class_involvedRewardGiveJob
logs_do_not_showroot cause of the Redis timeout, redis-server's own logs, infrastructure metrics (CPU/memory/network), any deploy or config change preceding the incident, who applied mitigation or when (no actor fields in logs)
graded checksfirst_error_ts · service_named · cascade_ordered · query · gaps_stated

Model's actual answer

show the model's full answer
First error

- 2026-09-03T14:01:12Z — `reward-service`
- Exact error: `Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s`
- No job ID or job class is attached to this first line.
- Earliest failure in `sidekiq_jobs.csv`: `J-00005` / `RewardGiveJob` at 14:01:46Z, 34 seconds later:
  `14:01:46 − 14:01:12 = 0m34s`

Cascade, in timestamp order

14:01:12  `reward-service` — Redis connection to `redis-primary:6379` times out.

14:01:20  `reward-service` — `RewardGiveJob` retry exhausted.

14:01:30  `reward-service` — `RewardGiveJob` retry exhausted.

14:01:40  `reward-service` — `RewardGiveJob` retry exhausted.

14:01:40  `sidekiq` — `RewardGiveJob` failed; retrying in 60s.

14:01:46  Job record — `J-00005`, `RewardGiveJob`, `Redis::TimeoutError`.

14:01:51  Job record — `J-00001`, `RewardGiveJob`, `Redis::TimeoutError`.

14:01:54  Job record — `J-00003`, `RewardGiveJob`, `Redis::TimeoutError`.

14:01:56  Job record — `J-00002`, `RewardGiveJob`, `Redis::TimeoutError`.

14:01:57  Job record — `J-00004`, `RewardGiveJob`, `Redis::TimeoutError`.

14:02:28  `sidekiq` — `RewardGiveJob` failed; retrying.

14:02:30  `sidekiq` — reward queue depth above 10,000.

14:02:36  Job record — `J-00013`, `RecognitionDigestJob`, `Redis::TimeoutError`.

14:02:51  Job records — `J-00007` and `J-00011`, `RewardGiveJob`, tied timestamp.

14:02:56  Job record — `J-00008`, `RewardGiveJob`.

14:02:57  Job records — `J-00010` and `J-00012`, `RewardGiveJob`, tied timestamp.

14:02:58  Job record — `J-00009`, `RewardGiveJob`.

14:03:05  `api-gateway` — `502 upstream timeout calling reward-service /gives`.

14:03:15  Job record — `J-00014`, `RecognitionDigestJob`.

14:03:30  `web-app` — Give form submission failed with an upstream 502 from `api-gateway`.

14:03:31  `sidekiq` — `RewardGiveJob` failed; retrying.

14:03:48  `api-gateway` — `502 upstream timeout calling reward-service`.

14:04:13  `api-gateway` — `502 upstream timeout calling reward-service`.

14:04:22  `sidekiq` — `RewardGiveJob` failed; retrying.

14:04:45  `web-app` — Give form submission failed with an upstream 502.

14:04:55  Job record — `J-00015`, `RecognitionDigestJob`.

14:05:16  `api-gateway` — `502 upstream timeout calling reward-service`.

14:05:26  `sidekiq` — `RewardGiveJob` failed; retrying.

14:05:42  `web-app` — Give form submission failed with an upstream 502.

14:05:50  Job record — `J-00016`, `RecognitionDigestJob`.

14:06:47  `sidekiq` — `RewardGiveJob` failed; retrying.

14:06:49  `web-app` — Give form submission failed with an upstream 502.

14:06:52  `api-gateway` — `502 upstream timeout calling reward-service`.

14:10:56–14:20:59  `postgres` — six `checkpoint complete` INFO events; no relationship to the Redis failure is shown.

14:22:10  `reward-service` — Redis connection restored; job processing resumed.

14:24:45  `sidekiq` — reward queue depth below 500.

Service and job map

- Originating service: `reward-service`
- Redis target: `redis-primary:6379`
- Primary failed job: `RewardGiveJob`
- Additional failed job class: `RecognitionDigestJob`
- Queue worker: `sidekiq`
- Request propagation: `api-gateway`, then `web-app`
- Job arithmetic: `12 RewardGiveJob + 4 RecognitionDigestJob = 16 failed job records`

Timing arithmetic

- First error to queue warning: `14:02:30 − 14:01:12 = 1m18s`
- First error to first API 502: `14:03:05 − 14:01:12 = 1m53s`
- First error to first form failure: `14:03:30 − 14:01:12 = 2m18s`
- First error to restored Redis: `14:22:10 − 14:01:12 = 20m58s`
- First error to queue below 500: `14:24:45 − 14:01:12 = 23m33s`

Datadog query to confirm the first error

Using the supplied CSV fields as Datadog search attributes:

`service:reward-service level:ERROR @timestamp:[2026-09-03T14:01:00Z TO 2026-09-03T14:01:20Z] @message:"Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s"`

What the logs do not show

- The underlying cause of the Redis timeout, such as Redis process failure, network failure, resource pressure, failover, or configuration change.
- Any remediation, deployment, restart, or scale action that restored the connection.
- The exact queue depth at any point, the peak backlog, enqueue rate, or drain rate.
- Which specific job ID, request ID, or trace corresponds to the 14:01:12 error or any API 502.
- Whether any of the 16 failed jobs subsequently completed successfully or whether rewards or recognition digests were delivered.
- How the `RecognitionDigestJob` failures appeared in `datadog_logs.csv`; only the job file records them.
- The meaning of `reward-service` saying “retry exhausted” while `sidekiq` subsequently says “retrying.”
- Any company, deal, or user aliases, request counts, affected-user counts, or business impact.
feature-flag-state-summaryPASS 1.00
Plain-English question: Summarize what each feature flag controls and who has it on.
Prompt — what the model saw
Summarize the state of the listed feature flags: what each controls per the code excerpt, which companies or segments have it on, targeting rules, and any flag with no code reference or no targets. Company counts must match the export. Do not describe a flag that is not in the export.
cost $0.0000 · 34s · in 37,997 / out 2,677 tok
✓ orphans✓ no_targets✓ counts_match✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: Recognition streaks on for 42 beta companies, redeem-flow redesign off with a 12-company target list, plus any flag with no code reference or no targets.
The answer in one line: rule = company counts must equal the export; do not describe flags absent from the export.
code_referencedrecognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, redeem_flow_redesign, analytics_dashboard_v3, ms_teams_app_v2
flags{"recognition_streaks_v2": {"state": "on", "targeting_rules": "segment:beta_companies", "company_count": 42}, "points_budget_guardrails": {"state": "on", "targeting_rules": "all_companies", "company_c
flags_with_no_targetspaused_offboard_cleanup
orphan_flags_no_code_referencelegacy_give_modal, survey_boosters_q3, paused_offboard_cleanup
rulecompany counts must equal the export; do not describe flags absent from the export.
graded checksorphans · no_targets · counts_match · no_invented

Model's actual answer

show the model's full answer
| Flag | State | What the code controls | Targeting rule | Company count |
|---|---|---|---|---:|
| `recognition_streaks_v2` | on | When enabled for a company, records recognition events through `StreakTracker.record(give)`. | `segment:beta_companies` | 42 |
| `points_budget_guardrails` | on | When enabled, enforces points via `BudgetService.new(company).enforce!(giver, points)`. | `all_companies` | 220 |
| `slack_dm_nudges` | on | When enabled, sends a Slack DM through `SlackDm.send_nudge(user)`. | `segment:region_na` | 87 |
| `redeem_flow_redesign` | off | When enabled, renders `RedeemV2Component`; otherwise renders `RedeemV1Component`. | `targeted_list` — companies not identified in the export | 12 |
| `analytics_dashboard_v3` | on | When enabled, creates `AnalyticsV3.new(company)`. | `segment:tier_three` | 65 |
| `ms_teams_app_v2` | off | When enabled, installs `TeamsAppV2` for the company. | `targeted_list` — companies not identified in the export | 9 |
| `legacy_give_modal` | off | No code reference in the excerpt. | `segment:legacy_plan` | 14 |
| `survey_boosters_q3` | on | No code reference in the excerpt. | `segment:legacy_plan` | 7 |
| `paused_offboard_cleanup` | off | No code reference in the excerpt. | No targeting rule supplied | 0 |

Exceptions:
- No code reference: `legacy_give_modal`, `survey_boosters_q3`, `paused_offboard_cleanup`.
- No targets: `paused_offboard_cleanup` only, with `company_count = 0`.
- `targeted_list` identifies the targeting method, but the export does not provide the individual companies.

Count arithmetic:
- On flags: 42 + 220 + 87 + 65 + 7 = 421.
- Off flags: 12 + 9 + 14 + 0 = 35.
- Sum of the export’s company-count field: 421 + 35 = 456.

These are counts by flag, not unique companies; the export does not indicate whether the same company appears under multiple flags.
nda-triagePASS 1.00
Plain-English question: Triage each NDA green, yellow, or red — and quote the clause that decides it.
Prompt — what the model saw
Triage each NDA as GREEN (standard approval), YELLOW (counsel review), or RED (full legal review). Cite the clause that drives each classification by quoting it. Catch embedded non-solicits, non-competes, missing carve-outs, and unusual term or governing law. One line of reasoning per NDA.
cost $0.0000 · 29s · in 18,736 / out 2,487 tok
✓ green✓ yellow✓ red✓ non_solicit_caught✓ clause_quoted✓ correct_labels

✓ Correct Answer (ground truth)

Plain-English answer: NDA-1 green (standard mutual terms), NDA-2 yellow for an embedded 18-month non-solicit, with the driving clause quoted for every classification.
The answer in one line: non_solicit_in = NDA-2
checksclassification_per_nda, clause_quotes_exist_in_document, non_solicit_caught, carveouts_checked
classifications{"NDA-1": {"classification": "GREEN", "driving_reason": "standard mutual NDA: 2-year term, standard carve-outs, Delaware law", "driving_clause_quote": "Carve-outs: information that (a) is or becomes p
non_solicit_inNDA-2
graded checksgreen · yellow · red · non_solicit_caught · clause_quoted · correct_labels

Model's actual answer

show the model's full answer
NDA-1 — [PARTY A] and Bonusly — GREEN — 2-year term + 3-year survival = 5 years if termination occurs at term end; mutual confidentiality, Delaware law, all five standard carve-outs, and “No license, no obligation to proceed, no exclusivity” make it standard: “Carve-outs: information that (a) is or becomes publicly available through no breach, (b) was known prior to disclosure, (c) is received from a third party without duty of confidence, (d) is independently developed, or (e) must be disclosed by law or court order.”

NDA-2 — [PARTY B] and Bonusly — YELLOW — the embedded non-solicit requires counsel review; 18 ÷ 12 = 1.5 years, and 3 + 1.5 = 4.5 years of restriction if termination occurs at term end: “During the term of this Agreement and for eighteen (18) months thereafter, neither party shall solicit for employment or hire any employee of the other party with whom it came into contact under this Agreement.”

NDA-3 — [PARTY C] and Bonusly — RED — full legal review required for a broad three-year non-compete, five-year one-way confidentiality term, zero stated carve-outs, and non-U.S. exclusive jurisdiction: “For a period of three (3) years following the Effective Date, Recipient shall not, directly or indirectly, engage in or provide services to any business that competes with Discloser's business”; “No confidentiality carve-outs are stated in this Agreement”; “This Agreement shall be governed by and construed in accordance with the laws of the Republic of Ireland, and the parties submit to the exclusive jurisdiction of its courts.”