← back to all models ← RevenueBench home
fireworks

fireworks/ember-1

Mean 0.980 · 36/40 perfect tests · $7.31 total · modalities in: text, image · out: text · each card: the prompt → the correct answer (gold) → this model's actual answer

What do the modalities mean?

Modalities in means what you can send this model: text, images, files, video, or audio. Modalities out means what it can send back. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what this model can handle in real work.

data-crm 5 tests
stage-hygiene-auditPASS 1.00
Plain-English question: How much of the open DS2+ pipeline is missing a substantive why-buys, and is anything missing DS1 timestamps or out of order?
Prompt — what the model saw
Audit stage hygiene on the open pipeline below. Rules: a deal at DS2 or later must have a substantive why-buys (15+ characters, not a placeholder such as tbd, n/a, see notes). A why_buys value of the form LONG(n) means the real text was present but truncated for display; treat it as present and substantive. Every deal should have a DS1 entry timestamp. A deal whose stage entry timestamps are out of order regressed (an earlier-stage timestamp is later than a later-stage timestamp). Report: percent of DS2+ deals non-compliant on why-buys overall and by owner (use the owners table for names), count of deals with no DS1 timestamp, count that regressed, and three process fixes. State definitions before numbers.
cost $0.2141 · 162s · in 14,196 / out 11,108 tok
✓ noncomp_pct✓ no_ds1_0✓ regressed_0✓ per_owner_breakdown✓ definitions_first✓ three_fixes

✓ Correct Answer (ground truth)

Plain-English answer: 81.5% of the 124 DS2+ deals are missing a substantive why-buys — Alex Franklin 48, Cole Ingram 19, Dana Mercer 18, Bryce Harmon 15, Farid Osman 1 — and every deal has a DS1 timestamp with none regressed.
The answer in one line: n_deals = 156; ds2plus = 124; noncomp_pct = 81.5; bryce_noncomp = 15; no_ds1 = 0; regressed = 0 (+1 more below)
n_deals156
bryce_noncomp15
ds2plus124
no_ds10
noncomp_pct81.5
regressed0
per_owner_noncompAlex Franklin=48 of 61 (78.7%); Cole Ingram=19 of 22 (86.4%); Dana Mercer=18 of 22 (81.8%); Bryce Harmon=15 of 18 (83.3%); Farid Osman=1 of 1 (100%)
graded checksnoncomp_pct · no_ds1_0 · regressed_0 · per_owner_breakdown · definitions_first · three_fixes

Model's actual answer

show the model's full answer
DEFINITIONS
- DS2+ = deals currently at stage DS2, DS3, DS4, or DS5 (DS1 excluded from why-buys rule).
- why-buys non-compliant = why_buys empty, <15 chars, or placeholder (tbd/n/a/see notes). LONG(n) counts as present and substantive per the brief.
- No-DS1 = t_ds1 empty.
- Regressed = any earlier-stage entry timestamp later than a later-stage entry timestamp (equal dates are not a regression).

SCOPE CHECK
- 156 open deals; 124 at DS2+; 32 at DS1.
- All populated why_buys values are LONG(n); zero placeholders found. All non-compliance is empty fields.
- All 156 deals have t_ds1. No timestamp inversions found.

1) WHY-BUYS NON-COMPLIANCE, DS2+ (empty why_buys)

Overall: 101 / 124 = 81.5% non-compliant (23 compliant, all LONG(n)).

By owner (names from owners.csv):
- Alex Franklin (84342457): 48 / 61 = 78.7%
- Dana Mercer (83155923): 18 / 22 = 81.8%
- Bryce Harmon (119337721): 15 / 18 = 83.3%
- Cole Ingram (83155924): 19 / 22 = 86.4%
- Farid Osman (716654662): 1 / 1 = 100% (Deal-499BF6)
- Elena Sinclair (701163055): 0 DS2+ deals, n/a
Check: 48+18+15+19+1 = 101; 61+22+18+22+1 = 124. ✓

2) DEALS WITH NO DS1 TIMESTAMP: 0

3) REGRESSED DEALS: 0
(Adjacent anomalies that are NOT regressions per the definition: 10 DS5 deals have no t_ds5; a few DS2/DS3 deals carry later-stage timestamps, e.g. Deal-0D0211, Deal-7E2131, Deal-BE3D9D at DS3 with t_ds4 set, and Deal-C7F9BF at DS2 with t_ds3 set — ordering within each is valid, but the stage history implies skipped/reverted stages.)

4) THREE PROCESS FIXES
a) Hard validation at stage advance: block transition to DS2+ unless why_buys is populated (15+ chars, placeholder values rejected). Today 81.5% of DS2+ deals — including COMMIT deals like Deal-D348E1, Deal-C26D20, Deal-547B2B, Deal-403845, Deal-B7EBD1, Deal-2465CE, Deal-A2B47C, Deal-C61CF7, Deal-FD9F4E — carry no why-buys at all, so forecast categories are uninspectable.
b) Owner-accountable backfill queue: push the 101-deal exception list into weekly forecast reviews, prioritized by forecast_category (COMMIT/BEST_CASE first) and close_date. Concentration is worst with Cole Ingram (86.4%) and Bryce Harmon (83.3%); Farid Osman's single DS2+ deal is non-compliant.
c) System-managed stage timestamps: make t_ds1..t_ds5 write-once, system-generated on transition, and require a timestamp for every stage up to current (all 10 DS5 deals lack t_ds5; several deals show later-stage stamps while sitting in an earlier stage). This removes backdating/ambiguity and makes regression and stage-skipping auditable.
owner-resolutionPASS 1.00
Plain-English question: Every open deal has a blank owner name — resolve each one to a real person using the owners table.
Prompt — what the model saw
The deal extract below has an owner name field that is blank on nearly every row and an owner id field that is populated. Using the owners table, resolve each open deal to an owner name. Report: number of deals resolved, deals whose owner id has no match in the owners table (list them), the owner ids that map to archived/deactivated owners (list their names), and the total pipeline amount per resolved owner. Do not guess a name for any unmatched id.
cost $0.1323 · 73s · in 17,755 / out 4,334 tok
✓ resolved_count✓ unmatched_none✓ deactivated_flagged✓ jake_amount✓ justin_amount✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: All 156 deals resolve to six named owners; jake_amount maps to Bryce Harmon ($1,054,144) and justin_amount to Alex Franklin ($624,310).
The answer in one line: jake_amount = 1,054,144; justin_amount = 624,310; resolved = 156
jake_amount1,054,144
justin_amount624,310
resolved156
deactivated
per_ownerBryce Harmon=1,054,144; Alex Franklin=624,310; Dana Mercer=341,195; Cole Ingram=288,161.43; Farid Osman=4,134; Elena Sinclair=2,100
graded checksresolved_count · unmatched_none · deactivated_flagged · jake_amount · justin_amount · no_invented

Model's actual answer

show the model's full answer
All 156 rows carry a populated owner_id, and every owner_id in the file appears in owners.csv. Join method: exact match on deals_open.owner_id = owners.owner_id.

**1) Deals resolved to an owner name: 156 of 156**

**2) Deals with no match in owners table: 0 — none to list**

**3) Owner ids mapping to archived/deactivated owners: none in the deal data**
The owners table flags two archived owners — Gavin Porter (1520255671) and Hugo Lindqvist (77260721) — but no deal in deals_open.csv references either id, so no open deal is owned by an archived owner.

**4) Total pipeline amount per resolved owner** (sum of `amount` over each owner's deals):

| Owner (owner_id) | Deals | Total pipeline |
|---|---|---|
| Bryce Harmon (119337721) | 35 | 1,054,144.00 |
| Alex Franklin (84342457) | 67 | 624,310.00 |
| Dana Mercer (83155923) | 24 | 341,195.00 |
| Cole Ingram (83155924) | 22 | 288,161.43 |
| Farid Osman (716654662) | 7 | 4,134.00 |
| Elena Sinclair (701163055) | 1 | 2,100.00 |

Arithmetic check: 1,054,144.00 + 624,310.00 + 341,195.00 + 288,161.43 + 4,134.00 + 2,100.00 = **2,314,044.43** total open pipeline across all 156 deals.

Notes on the data as given: three Bryce Harmon deals carry an amount of 1 (Deal-012CB1, Deal-483B2D, Deal-3795AD) and four Farid Osman deals carry an amount of 1 (Deal-03BA69, Deal-117863, Deal-F17780, Deal-2B39B0, Deal-8FDCD2 — five, not four) — these look like placeholder amounts and are included as-is since the task says to use only the data provided. Farid Osman's total of 4,134.00 = 1,249 (Deal-499BF6) + 2,880 (Deal-8BA24E) + five $1 placeholder deals.
gong-hubspot-join-integrityPASS 1.00
Plain-English question: How many call records point at deal ids that do not exist in the open pipeline?
Prompt — what the model saw
Check the integrity of the call-to-deal mapping below. The gong table lists calls in the last 90 days keyed to HubSpot deal ids; deals_open lists the current open default-pipeline deals. Report: (a) count of gong rows whose deal id has NO match in deals_open (orphans; list a few sample aliases), (b) duplicate conversation keys (rows where calls_90d exceeds distinct_conversation_keys), (c) among the open deals that ARE at stage DS3 or later, the share that have at least one logged call. Give sample aliases for the orphan finding, all taken from the data. Note: gong_owner_id is a Gong user id, not a HubSpot owner id.
cost $0.4310 · 249s · in 52,845 / out 16,836 tok
✓ orphan_count✓ dup_keys_0✓ ds3_coverage✓ sample_ids_present

✓ Correct Answer (ground truth)

Plain-English answer: 30 Gong call rows are orphans keyed to dead deal ids, none are duplicates, and DS3 call coverage is 29.4%.
The answer in one line: orphans = 30; ds3_cov_pct = 29.4; dup = 0
orphans30
ds3_cov_pct29.4
dup0
orphan_aliasesDeal-8FA85D, Deal-8FC3F9, Deal-3B7945, Deal-42B265, Deal-9CCC42, Deal-36EA09
graded checksorphan_count · dup_keys_0 · ds3_coverage · sample_ids_present

Model's actual answer

show the model's full answer
Call-to-deal mapping integrity check (gong_calls_by_deal_90d.csv vs deals_open.csv, joined on hs_deal_id = deal_id; gong_owner_id ignored per note):

(a) Orphan gong rows: 30 of 67
Gong rows whose deal id has no match in deals_open: 30/67 (44.8%). Arithmetic: 67 gong rows, 37 matched to the 156 open deals, 30 unmatched.
Sample orphan aliases (from the data):
- Deal-8FA85D (id 60251290957, 46 calls, last_call 2026-11-05)
- Deal-8FC3F9 (id 60251649055, 24 calls, last_call 2026-10-19)
- Deal-3B7945 (id 60251639682, 21 calls, last_call 2026-10-20)
- Deal-42B265 (id 61227242540, 21 calls, last_call 2026-11-25)
- Deal-9CCC42 (id 61430316324, 17 calls, last_call 2026-11-27)
Note: several orphans have last_call dates after 2026-09-24 (e.g. 2026-11-05, 2026-11-25, 2026-11-27), which is inconsistent with a "last 90 days" window as of today; flagging as a data-quality observation, not a separate finding.

(b) Duplicate conversation keys: 0
No row has calls_90d > distinct_conversation_keys. In all 67 rows, calls_90d == distinct_conversation_keys, so there are zero duplicate-key rows.

(c) DS3+ open deals with >= 1 logged call: 25/85 = 29.4%
Open deals at stage DS3 or later: 85 (DS3: 61, DS4: 14, DS5: 10).
Of these, 25 appear in the gong table. Share = 25/85 = 29.4%.
By stage: DS3 12/61 (19.7%), DS4 6/14 (42.9%), DS5 7/10 (70.0%).
snowflake-sql-trapped-catalogPASS 1.00
Plain-English question: Write one Snowflake query for first-month usage per customer using only the provided, trap-laden catalog.
Prompt — what the model saw
Using ONLY the catalog provided, write one Snowflake SQL query returning, per customer company, for its first calendar month as a customer: unique givers, recognition count, successful redemption count. Read the catalog notes carefully: some tables are stale or unpopulated and must be avoided, and one documented filter (excluding deleted givers) must NOT be applied to historical giving counts. Put a comment above each FROM or JOIN explaining the table choice. Return only fenced SQL.
cost $0.0404 · 43s · in 2,585 / out 1,849 tok
✓ avoids_stale_deals✓ avoids_unpopulated_live✓ uses_cohort_summary✓ succeeded_filter✓ no_giver_deletion_filter✓ returns_sql

✓ Correct Answer (ground truth)

Plain-English answer: A single correct SQL query that uses the right catalog tables and avoids the trap columns — graded on the exact tables and columns.
The answer in one line: use_table = PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
avoid_stalePRODUCTION.HUBSPOT.DEALS (stale)
avoid_unpopulatedV2_LIVE.OBJECTS_DEALS (unpopulated)
forbidden_filterdeleted-giver exclusion must NOT be applied
outputone fenced SELECT with per-FROM/JOIN table-choice comment
required_filtersucceeded redemptions only
use_tablePRODUCTION.PLG.COMPANY_COHORT_SUMMARY
graded checksavoids_stale_deals · avoids_unpopulated_live · uses_cohort_summary · succeeded_filter · no_giver_deletion_filter · returns_sql

Model's actual answer

show the model's full answer
```sql
-- Grain note: PRODUCTION.PLG.COMPANY_COHORT_SUMMARY is one row per self-serve company,
-- so no aggregation or joins are required; M1_* columns are the pre-computed
-- first-calendar-month metrics (no arithmetic needed beyond selecting them).

SELECT
    COMPANY_ID,                                  -- company key column name is not documented in the catalog excerpt; adjust to the actual key
    FIRST_SUB_PAYMENT_DATE          AS became_customer_date,
    DATE_TRUNC('month', FIRST_SUB_PAYMENT_DATE) AS first_customer_month,
    M1_USERS                        AS unique_givers,                -- CAVEAT: catalog documents no M1 unique-givers column; M1_USERS is the closest documented M1 people metric
    M1_GIVES                        AS recognition_count,            -- no deleted-giver exclusion applied, per the documented rule (it would understate history)
    M1_REDEMPTIONS                  AS successful_redemption_count   -- documented as redemptions; succeeded-only per the global redemption rule
-- FROM: PRODUCTION.PLG.COMPANY_COHORT_SUMMARY is the ONLY catalog table carrying
-- per-company first-month (M1) giving and redemption metrics. Chosen over:
--   - PRODUCTION.HUBSPOT.DEALS (STALE, last sync 2023-03 — must not be used),
--   - HUBSPOT_HUB_1973303.V2_LIVE.OBJECTS_DEALS (UNPOPULATED — must not be used),
--   - PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 (redemption events only;
--     has no giver/recognition data and is not finance-grade-confirmed, so the cohort
--     table's M1_REDEMPTIONS is preferred for a first-month-per-company view).
FROM PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
WHERE FIRST_SUB_PAYMENT_DATE IS NOT NULL  -- restrict to companies that actually became customers (first subscription payment made)
```
crm-completeness-enrichmentPASS 1.00
Plain-English question: Audit the CRM for missing and invalid fields, and propose the fixes.
Prompt — what the model saw
Audit the CRM extract for completeness and propose fixes. Deals need owner, stage, amount, close date, why-buys; companies need industry, employee count, HQ country; contacts need email, title, persona. Report completeness percent per field. Identify duplicate company clusters (name variants, shared domain) with a survivor per cluster. List invalid emails and domain mismatches. Fill missing company fields only where the enrichment export has a matching row; where CRM and enrichment disagree list both and recommend a source. Never invent a value. End with the 10 fixes with the most pipeline amount at stake.
cost $0.1320 · 89s · in 10,448 / out 5,533 tok
✓ dup_clusters✓ invalid_emails✓ survivor_language✓ completeness_pct✓ no_invent

✓ Correct Answer (ground truth)

Plain-English answer: 9 companies missing employee count, 6 missing HQ country, 4 invalid emails, 1 domain mismatch, 2 duplicate clusters, and 25 enrichment matches.
The answer in one line: n_contacts = 52; n_companies = 34; enrichment_matches = 25; invalid_email_count = 4; n_dup_clusters = 2; domain_mismatch_count = 1 (+4 more below)
n_contacts52
domain_mismatch_count1
enrichment_matches25
invalid_email_count4
n_companies34
n_dup_clusters2
company_missingindustry=0; employee_count=9; hq_country=6
domain_mismatchesCT-0011
duplicate_clusters['C-0A092931', 'C-0A092932'], ['C-0A092933', 'C-0A092934']
invalid_emailsCT-0010, CT-0080, CT-0081, CT-0192
graded checksdup_clusters · invalid_emails · survivor_language · completeness_pct · no_invent

Model's actual answer

show the model's full answer
MISSING INPUT, stated up front: no deals file was provided. Deal-field completeness (owner, stage, amount, close date, why-buys) cannot be audited, and the "top 10 fixes by pipeline amount" cannot be computed on real amounts — there is no amount data anywhere in the extract. I will not invent either. Where a fix below touches a company with pipeline, the amount is marked "unknown."

1. COMPLETENESS PER FIELD

Companies (n=34)
- domain: 34/34 = 100.0%
- industry: 34/34 = 100.0% (but see normalization issues below — populated ≠ clean)
- employee_count: 25/34 = 73.5% (9 missing)
- hq_country: 28/34 = 82.4% (6 missing)

Contacts (n=52)
- email: 52/52 = 100.0% populated, but 4/52 = 7.7% are invalid values ("user0@" etc.), so valid-email completeness is 48/52 = 92.3%
- title: 39/52 = 75.0% (13 missing)
- persona: 37/52 = 71.2% (15 missing)

Deals: not auditable — no file.

2. DUPLICATE COMPANY CLUSTERS (shared domain)

Cluster A — acme-corp.com: C-0A092931 (Technology, 500, US) + C-0A092932 (tech, 510, USA)
- Survivor: C-0A092931 (canonical casing, cleaner industry value; employee counts conflict 500 vs 510 — needs manual confirmation, enrichment has no acme-corp.com row so no tiebreaker available).

Cluster B — globex.io: C-0A092933 (SaaS, 200, US) + C-0A092934 (Technology, 200, US)
- Survivor: C-0A092934 (industry "Technology" matches the taxonomy used across the rest of the file; "SaaS" appears nowhere else). Employee count and country agree, so the only loss on merge is the "SaaS" label.

No name-variant clusters beyond these two — all other aliases are opaque IDs, so shared domain is the only detectable signal.

3. INVALID EMAILS (4)
- CT-0010 (C-66D1FC): "user0@" — no domain
- CT-0080 (C-92D97D): "user0@" — no domain
- CT-0081 (C-92D97D): "user1@" — no domain
- CT-0192 (C-425E2A): "user2@" — no domain

4. DOMAIN MISMATCHES (1)
- CT-0011 (C-66D1FC): user1@other-domain.com vs company domain 66d1fc.com. Either the email or the company association is wrong — flag for rep review, do not auto-correct.
- The contact.domain column matches company.domain on all 52 rows (no mismatches there).

5. COMPANY FIELD FILLS FROM ENRICHMENT (only where a matching ZI row has a value)

Safe fills (ZI row exists and has the value):
- C-EC3025 employee_count = 400
- C-96039F employee_count = 400
- C-44EA29 employee_count = 400
- C-D04904 employee_count = 400
- C-B23205 employee_count = 400
- C-60C75F employee_count = 400
- C-7BBDFA employee_count = 400
- C-50D386 employee_count = 400

Cannot fill (ZI row missing the field or no ZI row — do NOT invent):
- hq_country: C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5 (ZI blank), C-EE9FFB (no ZI row)
- employee_count: C-93C8BF (no ZI row)

6. CRM vs ENRICHMENT DISAGREEMENTS (recommendation per case)

Industry, 10 rows — CRM says tech/Technology, ZI says Computer Software:
C-66D1FC, C-EC3025, C-44EA29, C-92D97D, C-D04904, C-77A95A, C-AA8DDA, C-B25F40, C-60C75F, C-425E2A.
Recommendation: keep CRM "Technology" as the canonical value and treat ZI's "Computer Software" as the same concept in a different taxonomy (10/10 of ZI's tech-ish rows use it). The real CRM problem is casing/whitespace noise: "tech" (C-66D1FC, C-44EA29, C-60C75F), "Tech " with trailing space (C-425E2A, C-BA969B, C-93C8BF, C-C9BB20) — normalize all to "Technology". Do not overwrite with ZI's label or you split your own taxonomy.

employee_count: C-0A092931 (500) vs C-0A092932 (510) — intra-CRM conflict, ZI silent; resolve during merge of Cluster A.

hq_country: no true disagreements after normalizing US/USA/United States → US; ZI agrees with CRM everywhere both are populated. CRM normalization needed: "USA" (C-EC3025, C-96039F, C-E51FB7, C-425E2A, C-2D7423), "United States" (C-C6FE92, C-60C75F) → "US". Also normalize "health care" → "Healthcare" (C-7BBDFA, C-50D386).

7. TOP FIXES BY PIPELINE AT STAKE

Pipeline amount per company is unknowable — no deals file. Ranking below is by record-impact count; amounts are unknown for every line and should be attached once the deals extract is provided.

1. Merge Cluster A (acme-corp.com) → survivor C-0A092931; resolve 500 vs 510 employee conflict. Pipeline at stake: unknown.
2. Merge Cluster B (globex.io) → survivor C-0A092934. Pipeline at stake: unknown.
3. Fix 4 invalid emails CT-0010, CT-0080, CT-0081, CT-0192 — these contacts (incl. two champions) are unreachable as-is. Unknown.
4. Resolve CT-0011 domain mismatch (user1@other-domain.com vs 66d1fc.com) — champion on C-66D1FC. Unknown.
5. Fill 8 missing employee_counts from ZI (C-EC3025, C-96039F, C-44EA29, C-D04904, C-B23205, C-60C75F, C-7BBDFA, C-50D386). Unknown.
6. Backfill 15 missing contact personas (37/52 = 71.2% complete) — worst contact-field gap. Unknown.
7. Backfill 13 missing contact titles (75.0% complete). Unknown.
8. Normalize industry values: 7 rows of tech/Tech /health care noise → canonical Technology/Healthcare. Unknown.
9. Normalize hq_country: 7 rows of USA/United States → US. Unknown.
10. hq_country gaps with no enrichment source (C-2D1F1B, C-D73B89, C-44EA29, C-D04904, C-2C60E5, C-EE9FFB) — route to manual research; do not fill from nothing. Unknown.

To turn #10-style ranking into a real amount-weighted list, provide the deals extract (owner, stage, amount, close date, why-buys) and I'll redo the deal completeness audit and re-rank.
deal-intelligence 5 tests
closed-lost-classification0.80
Plain-English question: Why did we lose the last six months of deals — one primary category per deal, and whose side is the reason on?
Prompt — what the model saw
Classify every deal lost in the last 6 months below into one primary category: pricing, competitor, no decision, timing, product gap, champion left, or other. Also assign a side (Bonusly, buyer, unknown). Then summarize: category counts, the side split, the number of deals where the structured closed_lost_tag clearly disagrees with the free-text reason, and the two patterns most worth acting on. Base everything only on the tag and free-text provided.
cost $0.4276 · 309s · in 37,870 / out 20,482 tok
✓ total_90✗ competitor_top✓ category_counts_present✓ side_split✓ two_patterns

✓ Correct Answer (ground truth)

Plain-English answer: Every closed-lost deal gets exactly one category (pricing, competitor, no decision, timing, product gap, champion left, other) plus a side, per the computed classification.
The answer in one line: total_lost_6mo = 90; top_reason = competitor (26)
total_lost_6mo90
categoriespricing, competitor, no decision, timing, product gap, champion left, other
patternstwo most actionable, tag-vs-text disagreements counted
sideBonusly / buyer / unknown per deal
top_reasoncompetitor (26)
graded checkstotal_90 · competitor_top · category_counts_present · side_split · two_patterns

Model's actual answer

show the model's full answer
PER-DEAL CLASSIFICATION (90 deals, classified from tag + free-text only)

deal_id      alias        tag                                    -> category      side       note
63027745829  Deal-DB0AAC  Lost- Timing (1 year or more)          -> timing        buyer
63683330727  Deal-F7F635  Competitor                             -> competitor    buyer
63327490589  Deal-AC944F  MIA                                    -> no decision   unknown
63027809948  Deal-214060  MIA                                    -> no decision   unknown
49134744746  Deal-91A056  Lost- Timing (1 year or more)          -> timing        buyer
48988037529  Deal-29326C  Lost- Timing (1 year or more)          -> timing        buyer
64524670260  Deal-5DB9B0  Lost- Does not fit ICP                 -> other         buyer      spam
63836912221  Deal-831B7B  Lost- Timing (1 year or more)          -> timing        buyer
63680220945  Deal-F97C37  Competitor                             -> product gap   Bonusly    competitor won on diversified offerings
41554388661  Deal-13E9CF  Doing nothing/Not a priority/Cost      -> no decision   buyer      deprioritized, not budget
63222333276  Deal-39E25C  Lost- Timing (1 year or more)          -> timing        buyer
63291006863  Deal-7ED004  Lost- Budget/Price                     -> pricing       buyer
59275344824  Deal-21B045  MIA                                    -> no decision   unknown
58754552851  Deal-B3ABED  Lost- Timing (1 year or more)          -> timing        buyer      MIA but explicit 2028 window
62455767176  Deal-422BA6  Competitor                             -> competitor    buyer      ADP TotalSource partnership
61050677765  Deal-ED9AE7  Lost DM                                -> timing        buyer      text: "Timing, budget, authority" — timing first
61038826051  Deal-988493  MIA                                    -> no decision   unknown
63222778291  Deal-381C8C  Competitor                             -> competitor    unknown    no context beyond "not moving forward"
59418526836  Deal-F308CA  MIA                                    -> no decision   unknown
62750632013  Deal-F1E8A6  Competitor                             -> competitor    unknown    no context
60035957084  Deal-B6AC09  Lost- Timing (1 year or more)          -> timing        buyer
62750599045  Deal-70F704  Lost DM                                -> no decision   unknown    DISAGREE: text says MIA, no DM loss
61873010467  Deal-E6E80A  Lost- Timing (1 year or more)          -> timing        buyer
54322940958  Deal-B038F0  Lost- Timing (1 year or more)          -> timing        buyer
61625438845  Deal-4664E1  MIA                                    -> no decision   unknown
63222258948  Deal-175756  Lost- Timing (1 year or more)          -> timing        buyer
63717524046  Deal-E74A73  Doing nothing/Not a priority/Cost      -> no decision   buyer      testing manually first
63661381816  Deal-DDAB52  Competitor                             -> competitor    buyer      Rippl named
63514024330  Deal-ACE061  Competitor                             -> competitor    buyer      HeyTaco suspected
62852981522  Deal-BB78F3  Lost- Timing (1 year or more)          -> timing        buyer
60984778911  Deal-D48E0B  MIA                                    -> no decision   unknown
61054009677  Deal-15DA99  Lost- Timing (1 year or more)          -> timing        buyer
49530802588  Deal-F4AF5D  Lost- Timing (1 year or more)          -> timing        buyer
62115565909  Deal-79B7A1  Lost- Timing (1 year or more)          -> timing        buyer
62487728289  Deal-583ADB  MIA                                    -> no decision   unknown
63680238945  Deal-8E27DA  Feature Request                        -> competitor    buyer      DISAGREE: went with a swag provider, not a feature gap
63433935544  Deal-2D2F8D  Competitor                             -> competitor    buyer
60694374202  Deal-E0441F  MIA                                    -> no decision   Bonusly    stale deal inherited from departed rep
60897501515  Deal-7CB44D  MIA                                    -> no decision   unknown
60848492546  Deal-0F96AA  Competitor                             -> competitor    buyer
60355222018  Deal-1BCA50  Competitor                             -> competitor    buyer      stakeholder already down path with other vendor
61625560885  Deal-7CC678  Competitor                             -> competitor    unknown    nothing specific provided
59370037379  Deal-FAC17C  Lost DM                                -> no decision   buyer      no exec approval, but text shows no DM loss
61052858247  Deal-242273  Competitor                             -> product gap   Bonusly    lost on points-currency/onsite spend capability
56896716581  Deal-50E5D8  Doing nothing/Not a priority/Cost      -> no decision   buyer
62706569880  Deal-A2C349  Competitor                             -> competitor    buyer      Awardco named
59729560611  Deal-9F176A  Lost- Timing (1 year or more)          -> timing        buyer
61764780962  Deal-7B2236  Doing nothing/Not a priority/Cost      -> pricing       buyer      wants simpler/cheaper
57663815975  Deal-AFA56C  MIA                                    -> no decision   unknown
61129576246  Deal-C7156E  Competitor                             -> competitor    buyer
60866104098  Deal-C33D91  Lost- Budget/Price                     -> pricing       buyer
59086317965  Deal-9048EB  MIA                                    -> product gap   Bonusly    DISAGREE: text says bad fit, multiple feature gaps
60857702003  Deal-5E64CE  Doing nothing/Not a priority/Cost      -> pricing       buyer      Nectar exit fee is the blocker
61415737717  Deal-8A0992  Competitor                             -> competitor    buyer
63085142442  Deal-D0C698  Competitor                             -> competitor    buyer      Kudos named
56549284976  Deal-69CF3D  Lost- Timing (1 year or more)          -> timing        buyer
61507337022  Deal-ECBF89  Lost- Timing (1 year or more)          -> timing        buyer
57663820059  Deal-3618CC  Lost DM                                -> product gap   Bonusly    DISAGREE: text says "Wanted Surveys"
60548236897  Deal-EECC02  Competitor                             -> competitor    buyer
60896018951  Deal-5AD03E  Competitor                             -> product gap   Bonusly    DISAGREE: wanted defined budget access (capability ask)
62121718303  Deal-D1A623  Lost- Timing (1 year or more)          -> timing        buyer
63189310018  Deal-413C56  Doing nothing/Not a priority/Cost      -> no decision   buyer
60008683142  Deal-47F1A1  Competitor                             -> competitor    buyer      WorkTango named
54352704007  Deal-BF2A98  Competitor                             -> competitor    buyer      HiThrive named
62115549771  Deal-2A292B  Doing nothing/Not a priority/Cost      -> no decision   buyer      building internally
60868303272  Deal-D1AABF  MIA                                    -> no decision   unknown
60331562409  Deal-FEDBCB  Doing nothing/Not a priority/Cost      -> no decision   buyer
62622503749  Deal-1E7DA9  Competitor                             -> competitor    buyer
61625500700  Deal-2BBA21  MIA                                    -> no decision   unknown
62852981127  Deal-286F9C  Competitor                             -> competitor    buyer
62704591183  Deal-7FBAC6  Doing nothing/Not a priority/Cost      -> no decision   buyer
60008716662  Deal-369281  Competitor                             -> competitor    buyer      Paylocity native
61475258733  Deal-386F6E  MIA                                    -> no decision   unknown
61114491171  Deal-9FCD0D  Competitor                             -> competitor    buyer      Canadian company preference
55624236610  Deal-55867E  Lost- Timing (1 year or more)          -> timing        buyer
62853160058  Deal-DAFB82  Lost- Budget/Price                     -> pricing       buyer      no budget until 2028
59370028385  Deal-2FEDDB  Doing nothing/Not a priority/Cost      -> no decision   buyer
61024624821  Deal-64B19A  Competitor                             -> competitor    buyer      Motivosity suspected
60419904928  Deal-3F86A0  MIA                                    -> no decision   unknown
60675690108  Deal-096750  MIA                                    -> no decision   unknown
61055126627  Deal-F325A5  Lost DM                                -> champion left buyer      layoffs + leadership change
61432497792  Deal-ABD14C  Doing nothing/Not a priority/Cost      -> no decision   buyer
60868240474  Deal-79E61A  MIA                                    -> no decision   unknown
60551632419  Deal-8A119B  Lost- Budget/Price                     -> pricing       buyer
60644185922  Deal-AE7C4E  MIA                                    -> no decision   unknown
62115387928  Deal-DAB4F1  MIA                                    -> no decision   unknown
61432389647  Deal-B4B50F  MIA                                    -> no decision   unknown
60786197933  Deal-981AD4  Feature Request                        -> product gap   Bonusly    UI + not UK focused
62121470977  Deal-DC77FE  Competitor                             -> product gap   Bonusly    text says price not a factor; lost on customization
61129575303  Deal-5885B9  MIA                                    -> no decision   unknown

SUMMARY

Category counts (n=90, total $1,267,945):
  no decision     32  (35.6%)  $283,264
  competitor      23  (25.6%)  $330,915
  timing          20  (22.2%)  $262,191
  product gap      7  ( 7.8%)  $190,565
  pricing          6  ( 6.7%)  $175,810
  champion left    1  ( 1.1%)  $14,400
  other            1  ( 1.1%)  $10,800

Side split:
  buyer    58  (64.4%)
  unknown  24  (26.7%)
  Bonusly   8  ( 8.9%)

Tag/text disagreements: 5 deals
  Deal-70F704  tag "Lost DM"         -> text describes MIA/no contact (no decision)
  Deal-8E27DA  tag "Feature Request" -> text: went with a swag provider (competitor)
  Deal-9048EB  tag "MIA"             -> text: bad fit, multiple feature gaps (product gap)
  Deal-3618CC  tag "Lost DM"         -> text: "Wanted Surveys" (product gap)
  Deal-5AD03E  tag "Competitor"      -> text: wanted more defined budget access (product gap)

Two patterns most worth acting on:

1. No-decision is the largest category (32 deals, 35.6%) and it is mostly a hygiene/engagement
   black hole, not a verified buyer choice. 17 of the 32 are pure "MIA/unresponsive/no response"
   with zero stated reason, so side is unknown for 24 of 32. These deals were closed with no
   learning attached. Fix: require a minimum evidence standard (e.g., a named next step or a
   disqualification reason) before a deal can be marked closed-lost as MIA — otherwise the
   biggest loss bucket stays unactionable.

2. A large share of "timing" losses are re-engageable, and competitor losses increasingly hide
   product gaps. 12 of 20 timing deals name an explicit future window (2027, 2028, "new year,"
   "circle back") — that is $262K of pipeline that should have a dated re-engagement task, not
   just a closed stage. Separately, 3 of the 5 tag/text disagreements are deals tagged Competitor
   or Lost DM whose text actually cites a missing capability (surveys, points-currency/onsite
   spend, budget access, customization) — meaning product-gap losses are undercounted in the
   structured field and the true product-gap signal is bigger than the 7 tagged deals suggest.
pipeline-tieringPASS 1.00
Plain-English question: Score all 156 open deals into exactly one of LOCK / ACTION / BUILD / REVIVE / WATCH / RISKY.
Prompt — what the model saw
Score the open pipeline below on these signals: forecast category, stage, meetings_30d, contact count, and engagement recency. Assign exactly one tier per deal: LOCK, ACTION, BUILD, REVIVE, WATCH, or RISKY (RISKY = the forecast category disagrees with the engagement evidence). Never tier a deal LOCK if it has zero meetings_30d. Do NOT list every deal. Return ONLY a compact JSON object with exactly these keys: - "tier_counts": object mapping each tier to the number of deals in it (must sum to the total deal count), - "tier_examples": object mapping each tier to up to 3 example deal aliases from the data, - "risky_deals": array of the deal aliases you tiered RISKY, - "lock_violations": number of deals tiered LOCK that have zero meetings_30d (should be 0), - "pipeline_shape": one short paragraph on the overall pipeline shape. Note: inbound_emails_30d is 0 for every row due to a data defect, so treat meetings_30d as the inbound signal.
cost $0.6017 · 481s · in 55,599 / out 26,268 tok
✓ json_returned✓ tier_vocabulary✓ tier_counts_sum✓ risky_deals_listed✓ no_lock_violation_claimed

✓ Correct Answer (ground truth)

Plain-English answer: All 156 open deals tiered exactly once across the six tiers using forecast category, stage, meetings, contacts, and recency.
The answer in one line: total = 156
total156
graded checksjson_returned · tier_vocabulary · tier_counts_sum · risky_deals_listed · no_lock_violation_claimed

Model's actual answer

show the model's full answer
{
  "tier_counts": {
    "LOCK": 2,
    "ACTION": 27,
    "BUILD": 7,
    "REVIVE": 43,
    "WATCH": 61,
    "RISKY": 16
  },
  "tier_examples": {
    "LOCK": ["Deal-C26D20", "Deal-944310"],
    "ACTION": ["Deal-25F752", "Deal-D348E1", "Deal-403845"],
    "BUILD": ["Deal-A5E80A", "Deal-483B2D", "Deal-BB8880"],
    "REVIVE": ["Deal-B7EBD1", "Deal-0660B4", "Deal-BA571A"],
    "WATCH": ["Deal-E53952", "Deal-5408B0", "Deal-9AAE5F"],
    "RISKY": ["Deal-523604", "Deal-3795AD", "Deal-01E193"]
  },
  "risky_deals": ["Deal-523604", "Deal-3795AD", "Deal-01E193", "Deal-C1FA6D", "Deal-7BBDFA", "Deal-93C8BF", "Deal-1CCE5C", "Deal-927338", "Deal-DCA846", "Deal-819506", "Deal-03BA69", "Deal-117863", "Deal-F17780", "Deal-8BA24E", "Deal-2B39B0", "Deal-8FDCD2"],
  "lock_violations": 0,
  "pipeline_shape": "Pipeline of 156 open deals is heavily early-stage: 32 DS1, 39 DS2, 61 DS3, 14 DS4, 10 DS5. Forecast mix: 11 COMMIT, 40 BEST_CASE, 105 PIPELINE. Only 24 deals are late-stage (DS4/DS5), and 101 of 156 deals recorded zero meetings in the last 30 days, leaving a thin top-of-funnel with limited near-term coverage."
}

Pipeline shape: 156 open deals skew early-stage (32 DS1 / 39 DS2 / 61 DS3 vs only 14 DS4 / 10 DS5) and forecast is 105 PIPELINE vs 11 COMMIT + 40 BEST_CASE; 101 of 156 deals had zero meetings_30d, so near-term coverage rests on a narrow late-stage base while the 16 RISKY deals (all forecast PIPELINE but showing strong meeting-driven engagement) represent understated upside and the 43 REVIVE deals represent a large dormant tail.
call-transcript-extractionPASS 1.00
Plain-English question: Extract the CRM write-back fields from each call transcript as JSON.
Prompt — what the model saw
For each transcript, extract CRM write-back fields as JSON: why-buys (prospect statements only), pain points, stakeholders from the speaker list, budget signal (prospect-stated or null), timeline signal, competitor mentioned (only if the prospect raised it), next step (only if explicitly agreed), objections, confidence. Never fill a field from what the rep said or hoped.
cost $0.0613 · 60s · in 3,428 / out 3,074 tok
✓ json_emitted✓ budget_signal✓ null_discipline✓ stakeholders✓ competitor_prospect_only

✓ Correct Answer (ground truth)

Plain-English answer: Per transcript: prospect-only why-buys, pain points, stakeholders, budget and timeline signals, and competitors — e.g. TX-001: '$40k earmarked', live before open enrollment.
The answer in one line: rep = Alex Franklin
allowed_newTX-001, TX-002, TX-003, TX-004, TX-005, TX-006, Achievers, Bucketlist, Nectar
per_transcript{"TX-001": {"why_buys": ["automating anniversary and birthday awards"], "pain_points": ["HR team of three cannot keep up with it manually", "track everything in a spreadsheet, and people slip through
repAlex Franklin
rep_sourced_nulls{"TX-003": ["budget_signal \u2014 only the rep mentioned pricing/budget"], "TX-004": ["next_step \u2014 rep proposed, prospect said 'no promises'"], "TX-006": ["budget_signal \u2014 rep decoy, prospec
graded checksjson_emitted · budget_signal · null_discipline · stakeholders · competitor_prospect_only

Model's actual answer

show the model's full answer
```json
[
  {
    "transcript_id": "TX-001",
    "deal_alias": "Deal-CFE7F4",
    "why_buys": [
      "Automating anniversary and birthday awards (VP People)"
    ],
    "pain_points": [
      "HR team of three cannot keep up with manual awards (VP People)",
      "Everything tracked in a spreadsheet; people slip through the cracks (HR Admin)"
    ],
    "stakeholders": ["VP People", "HR Admin"],
    "budget_signal": "~$40k earmarked for engagement tools this fiscal year (VP People)",
    "timeline_signal": "Live before open enrollment in November (VP People)",
    "competitor_mentioned": "Achievers — prospect looked at it last year; too heavy for a team their size",
    "next_step": "Security review with IT lead on September 12 (explicitly agreed)",
    "objections": [
      "Need SSO and audit logs for IT to sign off (HR Admin)"
    ],
    "confidence": "high"
  },
  {
    "transcript_id": "TX-002",
    "deal_alias": "Deal-70BB30",
    "why_buys": [
      "Tie recognition to retention for the hourly workforce (Head of Total Rewards)"
    ],
    "pain_points": [
      "Regretted turnover over 30% in the hourly workforce (Head of Total Rewards)"
    ],
    "stakeholders": ["Head of Total Rewards", "CFO"],
    "budget_signal": "$25k pilot budget approved by finance for this quarter (CFO)",
    "timeline_signal": "Decision by end of September (CFO)",
    "competitor_mentioned": null,
    "next_step": "Send pilot agreement; prospect will route it to legal this week (explicitly agreed)",
    "objections": [
      "Workday integration must be rock solid — CFO's one condition"
    ],
    "confidence": "high",
    "note": "Prospect stated this is the first vendor they've had a real demo with — no competitor raised."
  },
  {
    "transcript_id": "TX-003",
    "deal_alias": "Deal-530B50",
    "why_buys": [
      "Make recognition visible across 12 retail locations (People Ops Manager)"
    ],
    "pain_points": [
      "Store managers have zero budget autonomy for on-the-spot recognition (People Ops Manager)"
    ],
    "stakeholders": ["People Ops Manager"],
    "budget_signal": null,
    "timeline_signal": "No rush until Q1 (People Ops Manager)",
    "competitor_mentioned": "Bucketlist — CEO used it at her last company and liked it",
    "next_step": "Schedule a call with the CEO; People Ops Manager will send two times (explicitly agreed)",
    "objections": [
      "CEO has to be sold first — she decides anything people-related",
      "CEO already likes Bucketlist"
    ],
    "confidence": "medium",
    "note": "CEO is referenced as the decision-maker but is not in the speaker list, so not included in stakeholders. The $8/employee/month figure was rep-stated, so budget_signal is null."
  },
  {
    "transcript_id": "TX-004",
    "deal_alias": "Deal-180D02",
    "why_buys": [
      "Consolidate three separate recognition tools into one (VP People)"
    ],
    "pain_points": [
      "Paying for three tools; none of them talk to the HRIS (VP People)"
    ],
    "stakeholders": ["VP People", "IT Security Lead"],
    "budget_signal": "Under $15k annually, VP People can approve without going to the board",
    "timeline_signal": "Procurement cycle runs six to eight weeks minimum (IT Security Lead)",
    "competitor_mentioned": null,
    "next_step": null,
    "objections": [
      "Security review took three months for the last vendor — IT Security Lead's hesitation",
      "Procurement cycle is six to eight weeks minimum"
    ],
    "confidence": "medium",
    "note": "CFO follow-up was proposed by the rep but not agreed ('Maybe — I need to check her calendar, no promises'), so next_step is null."
  },
  {
    "transcript_id": "TX-005",
    "deal_alias": "Deal-F8767A",
    "why_buys": [
      "Automate service milestones (HR Director)",
      "Analytics on recognition equity across departments (HR Director)"
    ],
    "pain_points": [
      "Night-shift teams feel invisible; engagement scores run 20 points lower (People Ops Coordinator)"
    ],
    "stakeholders": ["HR Director", "People Ops Coordinator"],
    "budget_signal": "$12k approved under the engagement line (HR Director)",
    "timeline_signal": "Running before the January all-hands (HR Director)",
    "competitor_mentioned": "Nectar — prospect is mid-pilot with Nectar; we'd need to beat that experience",
    "next_step": "Present to the prospect's exec team on October 2 (explicitly agreed)",
    "objections": [
      "Exec team is skeptical after a failed rollout two years ago",
      "Must beat the incumbent Nectar pilot experience"
    ],
    "confidence": "high"
  },
  {
    "transcript_id": "TX-006",
    "deal_alias": "Deal-EE195F",
    "why_buys": [
      "Cut the admin time on service awards (HR Manager)"
    ],
    "pain_points": [
      "HR Manager personally spends five hours a month ordering and shipping plaques"
    ],
    "stakeholders": ["HR Manager"],
    "budget_signal": null,
    "timeline_signal": "Fiscal year resets in February; a Q1 start is realistic (HR Manager)",
    "competitor_mentioned": null,
    "next_step": "Send the one-page overview; HR Manager will forward it to the COO this week (explicitly agreed)",
    "objections": [
      "COO usually prefers building things in-house"
    ],
    "confidence": "medium",
    "note": "Prospect stated 'Budget isn't the issue — time is,' so no budget figure exists; recorded as null rather than invented. No vendor competitor raised — the alternative is an internal build. COO is referenced but not in the speaker list, so not listed as a stakeholder."
  }
]
```

Arithmetic note: no derived calculations were required; all figures ($40k, $25k, $15k, $12k, 30%, 20 points, 5 hours, 12 locations, 6–8 weeks, 3 months) are restated verbatim from prospect lines. TX-002's ">30%" and TX-005's "20 points lower" are single prospect-stated values with no second data point to compute against.
next-to-close0.75
Plain-English question: Which 3 deals are closest to signature right now?
Prompt — what the model saw
Which 3 deals are closest to signature right now? Use the CRM data and the Slack thread; Slack often has the real status before the CRM does. For each: deal, amount, why it is close, what is left. Under 250 words.
cost $0.0272 · 24s · in 2,752 / out 934 tok
✗ exactly_three✓ top3_named✓ slack_first✓ blocked_excluded

✓ Correct Answer (ground truth)

Plain-English answer: Deal-547B2B ($11,200), Deal-B7EBD1 ($9,000), and Deal-403845 ($9,000) — all DS5 COMMIT deals closing September 10-11, with what is left on each.
blocked_excludeddeal_alias=Deal-2465CE; reason=champion left + procurement freeze (Slack #sales-sparring 2026-09-04)
exactly_three1
ranking_ruleSlack signature-imminent deal ranks first; remaining ranked by stage (DS5>DS4), forecast (COMMIT>BEST_CASE>PIPELINE), close_date asc, amount desc
top3rank=1; deal_alias=Deal-547B2B; amount=11200; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11 | rank=2; deal_alias=Deal-B7EBD1; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-10 | rank=3; deal_alias=Deal-403845; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11
graded checksexactly_three · top3_named · slack_first · blocked_excluded

Model's actual answer

show the model's full answer
Top 3 closest to signature (Slack overrides CRM where they conflict):

1. Deal-547B2B — $11,200 (Alex Franklin, DS5/COMMIT, close 2026-09-11)
Why close: Slack (09-04) says redlines came back clean, signing page is out, VP People committed to signing "tomorrow" — i.e., 2026-09-05, ahead of the CRM close date.
Left: nothing but the signature itself.

2. Deal-403845 — $9,000 (Alex Franklin, DS5/COMMIT, close 2026-09-11)
Why close: DS5/COMMIT in CRM, and Slack confirms it's "moving fine" — order form is already with their finance team. That's the last internal step before signature.
Left: finance approval, then signature.

3. Deal-A2B47C — $6,360 (Alex Franklin, DS5/COMMIT, close 2026-09-11)
Why close: DS5/COMMIT with a 09-11 close date, and Slack confirms it's "still warm, just normal legal-review pace" — no blockers flagged.
Left: standard legal review completion.

Explicitly excluded despite CRM signals:
- Deal-2465CE ($5,400, DS5/COMMIT, 09-10): CRM looks hot, but Slack says the champion left, procurement froze new vendors, and Dana is pulling it from commit — realistically Q4. Not close.
- Deal-B7EBD1 ($9,000, DS5/COMMIT, 09-10): earliest close date in CRM, but zero Slack corroboration; ranked below the three deals with confirmed live status.
- Deal-D348E1 ($13,770, DS5/COMMIT): warm per Slack, but close date is 2026-10-15 — healthy, not imminent.

Combined value of the top 3: $11,200 + $9,000 + $6,360 = $26,560.
product-gap-detectionPASS 1.00
Plain-English question: Which prospect complaints are true product gaps versus rollout or enablement issues?
Prompt — what the model saw
Surface true product gaps raised by prospects in active deals. For each candidate: quote the prospect line, classify as real gap, plan gate, or rollout/enablement issue using the product docs excerpt, and name the deal and amount. Only real gaps go in the summary. Rep language does not count as prospect voice.
cost $0.0340 · 30s · in 2,359 / out 1,469 tok
✓ real_gaps✓ plan_gate_classified✓ enablement_classified✓ gap_summary

✓ Correct Answer (ground truth)

Plain-English answer: Two real gaps (SCIM provisioning, ADP Workforce Now integration) separated from plan gates and enablement issues, each with the prospect quote and deal.
allowed_newTG-001, TG-002, TG-003, TG-004, TG-005, SCIM user provisioning, ADP Workforce Now integration, Custom report builder
enablement_issuestranscript_id=TG-004; deal_alias=Deal-84DBA6; amount=16000; feature=Slack integration recognition adoption; quote=We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.; why_enablement=Slack integration exists on all plans; managers were never trained
excluded_rep_voicetranscript_id=TG-005; line=Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap.; reason=rep voice, not prospect-raised
plan_gatestranscript_id=TG-003; deal_alias=Deal-CFE7F4; amount=32000; feature=Custom report builder; tier_required=Enterprise; quote=I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?
real_gapstranscript_id=TG-001; deal_alias=Deal-EC3025; amount=62000; feature=SCIM user provisioning; quote=We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.; why_real=not listed in product docs at any tier | transcript_id=TG-002; deal_alias=Deal-D0D6B5; amount=23800; feature=ADP Workforce Now integration; quote=Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.; why_real=supported HRIS list is Workday/BambooHR/Gusto/Rippling only
graded checksreal_gaps · plan_gate_classified · enablement_classified · gap_summary

Model's actual answer

show the model's full answer
GAP REVIEW — 5 transcripts, 5 candidate lines (4 prospect-voiced, 1 rep-voiced)

1) TG-001 / Deal-EC3025 — REAL GAP
Quote: "We need SCIM user provisioning — our security team will not approve any tool that requires manual account management." (Prospect, IT Security Lead)
Basis: product_docs.md states "SCIM user provisioning ... [is] NOT currently listed as supported capabilities." Not in any tier. Prospect-voiced, blocking requirement.

2) TG-002 / Deal-D0D6B5 — REAL GAP
Quote: "Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us." (Prospect, HRIS Manager)
Basis: supported HRIS list is Workday, BambooHR, Gusto, Rippling only; docs explicitly state "ADP Workforce Now integrations are NOT currently listed as supported capabilities." Prospect-voiced, self-declared dealbreaker.

3) TG-003 / Deal-CFE7F4 — PLAN GATE (excluded from summary)
Quote: "I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?" (Prospect, VP People)
Basis: Custom report builder exists in the product, gated to Enterprise tier per the plans table; rep confirmed same. Capability exists — this is packaging, not a missing feature.

4) TG-004 / Deal-84DBA6 — ROLLOUT/ENABLEMENT ISSUE (excluded from summary)
Quote: "We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it." (Prospect, People Ops Manager)
Basis: Slack integration is supported on all plans. Feature is live; the failure is adoption/training, not product capability.

5) TG-005 / Deal-36C33F — EXCLUDED, NOT PROSPECT VOICE
The mobile-app-for-hourly-workers statement came from Alex Franklin (rep): "we don't have a native mobile app for hourly workers yet." Rep language does not count per your rules. The prospect's only line ("The web version should be fine for our office staff for now") accepts the status quo and raises no requirement.

DEAL AMOUNTS
No amount field exists in transcripts_gaps.csv and no amounts appear in product_docs.md. Amounts for Deal-EC3025 and Deal-D0D6B5 (and all deals) are missing from the data provided — I can't state them without inventing numbers.

SUMMARY — TRUE PRODUCT GAPS (prospect-voiced, active deals)
Arithmetic: 5 candidate lines − 1 rep-voiced (TG-005) − 1 plan gate (TG-003) − 1 rollout issue (TG-004) = 2 real gaps.

- SCIM user provisioning — Deal-EC3025 (amount: not in data). Security blocker: "our security team will not approve any tool that requires manual account management."
- ADP Workforce Now integration — Deal-D0D6B5 (amount: not in data). Dealbreaker: "that's a dealbreaker for us."

Both gaps are corroborated by the product docs excerpt, which explicitly lists neither as supported.
rep-performance 5 tests
stale-pipeline-by-rep0.83
Plain-English question: Which open deals have had no email, call, or meeting in the last 7 days?
Prompt — what the model saw
The snapshot date is 2026-09-05. Using the open deals and the engagements-by-deal table, list every open deal with no logged email, call, or meeting in the last 7 days. Compute recency from the engagements table's last_email/last_call/last_meeting fields, NOT from any date field on the deal (the deal's last_contacted_field is unreliable). For each stale deal give: deal alias, owner name, stage, amount, days since last contact. Group by owner, order by amount descending within each owner. Then state, per owner, how many stale deals they carry and the total stale amount. (Do not write outreach text.)
cost $0.8911 · 532s · in 120,740 / out 27,927 tok
✓ stale_recall_ge_80✓ grouped_by_owner✓ jake_top_amount✓ kevin_stale_amt✓ jenn_stale_amt✗ stale_count

✓ Correct Answer (ground truth)

Plain-English answer: 65 open deals are stale; the largest is Deal-2D1F1B at $240,000, with Bryce Harmon carrying the most stale dollars at $626,243.
The answer in one line: jake_amt = 626,243; jenn_amt = 261,645; kevin_amt = 252,905.03; top_amount = 240,000; stale_count = 65; top_alias = Deal-2D1F1B
jake_amt626,243
jenn_amt261,645
kevin_amt252,905.03
stale_count65
top_amount240,000
top_aliasDeal-2D1F1B
graded checksstale_recall_ge_80 · grouped_by_owner · jake_top_amount · kevin_stale_amt · jenn_stale_amt · stale_count

Model's actual answer

show the model's full answer
All open deals with no logged email, call, or meeting in the 7 days ending 2026-09-05 (window 2026-08-29 through 2026-09-05; recency = max of last_email / last_call / last_meeting from engagements_by_deal_90d.csv; last_contacted_field not used).

Two open deals have no row in the engagements table at all, so recency cannot be computed for them; excluded from the stale list and flagged here: Deal-3EED2C (Alex Franklin, DS2, $7,200) and Deal-57FF13 (Elena Sinclair, DS1, $2,100).

63 stale deals, $1,243,129.03 total.

Bryce Harmon — 13 stale deals, $626,243.00
  Deal-2D1F1B   DS1   $240,000.00   81 days (last 2026-06-16)
  Deal-66D1FC   DS1   $ 99,000.00   16 days (last 2026-08-20)
  Deal-950043   DS1   $ 70,000.00   19 days (last 2026-08-17)
  Deal-B23205   DS1   $ 45,000.00   16 days (last 2026-08-20)
  Deal-7BBDFA   DS3   $ 37,440.00   46 days (last 2026-07-21)
  Deal-332637   DS2   $ 36,000.00    9 days (last 2026-08-27)
  Deal-1BEEBF   DS1   $ 31,500.00   19 days (last 2026-08-17)
  Deal-C5658B   DS1   $ 23,400.00   16 days (last 2026-08-20)
  Deal-40522D   DS3   $ 21,000.00   19 days (last 2026-08-17)
  Deal-F0EBBB   DS3   $ 11,400.00   24 days (last 2026-08-12)
  Deal-E25A09   DS1   $  6,000.00    9 days (last 2026-08-27)
  Deal-C9C286   DS2   $  5,502.00    9 days (last 2026-08-27)
  Deal-012CB1   DS1   $      1.00   23 days (last 2026-08-13)
  Sum check: 240000+99000+70000+45000+37440+36000+31500+23400+21000+11400+6000+5502+1 = 626,243

Alex Franklin — 18 stale deals, $102,336.00
  Deal-CC08D1   DS1   $ 24,000.00   16 days (last 2026-08-20)
  Deal-E73427   DS3   $ 18,000.00   10 days (last 2026-08-26)
  Deal-885F45   DS2   $  9,300.00   12 days (last 2026-08-24)
  Deal-C2FF3C   DS1   $  8,316.00   10 days (last 2026-08-26)
  Deal-0D2F7A   DS3   $  5,100.00   12 days (last 2026-08-24)
  Deal-6C60D4   DS3   $  4,800.00   12 days (last 2026-08-24)
  Deal-13FEBD   DS2   $  4,680.00   12 days (last 2026-08-24)
  Deal-9D0060   DS3   $  3,840.00   12 days (last 2026-08-24)
  Deal-690476   DS2   $  3,600.00   18 days (last 2026-08-18)
  Deal-C6D97A   DS4   $  3,240.00    8 days (last 2026-08-28)
  Deal-EE195F   DS3   $  3,120.00    8 days (last 2026-08-28)
  Deal-278DEC   DS3   $  2,700.00    8 days (last 2026-08-28)
  Deal-635B8E   DS3   $  2,600.00   18 days (last 2026-08-18)
  Deal-6883F3   DS1   $  2,400.00   16 days (last 2026-08-20)
  Deal-4A13AD   DS3   $  2,160.00   26 days (last 2026-08-10)
  Deal-F67D31   DS2   $  1,800.00    8 days (last 2026-08-28)
  Deal-5FDCE4   DS3   $  1,600.00   12 days (last 2026-08-24)
  Deal-BA571A   DS4   $  1,080.00   18 days (last 2026-08-18)
  Sum check: 24000+18000+9300+8316+5100+4800+4680+3840+3600+3240+3120+2700+2600+2400+2160+1800+1600+1080 = 102,336

Dana Mercer — 14 stale deals, $261,645.00
  Deal-44EA29   DS2   $ 60,000.00   10 days (last 2026-08-26)
  Deal-E51FB7   DS2   $ 43,875.00   12 days (last 2026-08-24)
  Deal-B42F46   DS1   $ 27,000.00   19 days (last 2026-08-17)
  Deal-BA3DDC   DS3   $ 23,400.00   15 days (last 2026-08-21)
  Deal-9DDE86   DS2   $ 20,000.00   15 days (last 2026-08-21)
  Deal-215CCA   DS3   $ 18,900.00   17 days (last 2026-08-19)
  Deal-5EED42   DS3   $ 16,250.00   11 days (last 2026-08-25)
  Deal-57887A   DS2   $ 15,000.00    8 days (last 2026-08-28)
  Deal-B7EBD1   DS5   $  9,000.00   16 days (last 2026-08-20)
  Deal-3974EB   DS4   $  9,000.00    8 days (last 2026-08-28)
  Deal-F40F04   DS2   $  8,100.00   15 days (last 2026-08-21)
  Deal-87DDD1   DS1   $  5,000.00   19 days (last 2026-08-17)
  Deal-F336B6   DS3   $  4,200.00   15 days (last 2026-08-21)
  Deal-0660B4   DS4   $  1,920.00   16 days (last 2026-08-20)
  Sum check: 60000+43875+27000+23400+20000+18900+16250+15000+9000+9000+8100+5000+4200+1920 = 261,645

Cole Ingram — 18 stale deals, $252,905.03
  Deal-D04904   DS2   $ 58,529.25   11 days (last 2026-08-25)
  Deal-B25F40   DS3   $ 40,000.00    8 days (last 2026-08-28)
  Deal-813836   DS2   $ 32,175.00   11 days (last 2026-08-25)
  Deal-1BA595   DS2   $ 31,750.00   11 days (last 2026-08-25)
  Deal-CFE1E8   DS3   $ 18,000.00   11 days (last 2026-08-25)
  Deal-CD47A6   DS2   $ 12,168.00   11 days (last 2026-08-25)
  Deal-627646   DS3   $ 11,193.00   11 days (last 2026-08-25)
  Deal-FF809F   DS2   $  7,781.20   11 days (last 2026-08-25)
  Deal-AF932D   DS2   $  7,225.40   11 days (last 2026-08-25)
  Deal-A71728   DS2   $  6,947.50   11 days (last 2026-08-25)
  Deal-8BC9F5   DS2   $  5,616.00   10 days (last 2026-08-26)
  Deal-175395   DS3   $  4,779.88   11 days (last 2026-08-25)
  Deal-481E24   DS3   $  4,140.00   10 days (last 2026-08-26)
  Deal-C7F9BF   DS2   $  3,360.00   11 days (last 2026-08-25)
  Deal-2F3A66   DS3   $  3,334.80   11 days (last 2026-08-25)
  Deal-342E96   DS2   $  2,700.00   24 days (last 2026-08-12)
  Deal-E568D5   DS3   $  1,875.00   11 days (last 2026-08-25)
  Deal-FD9F4E   DS5   $  1,330.00   10 days (last 2026-08-26)
  Sum check: 58529.25+40000+32175+31750+18000+12168+11193+7781.20+7225.40+6947.50+5616+4779.88+4140+3360+3334.80+2700+1875+1330 = 252,905.03

Farid Osman — 0 stale deals.
Elena Sinclair — 0 stale deals (her one open deal, Deal-57FF13, has no engagement data; see note above).

Per-owner summary:
  Bryce Harmon:   13 stale deals, $626,243.00
  Dana Mercer:    14 stale deals, $261,645.00
  Cole Ingram:    18 stale deals, $252,905.03
  Alex Franklin:  18 stale deals, $102,336.00
  Farid Osman:     0
  Elena Sinclair:  0
  Grand total:    63 stale deals, $1,243,129.03  (626,243.00 + 261,645.00 + 252,905.03 + 102,336.00 = 1,243,129.03)

Notes on method:
- "Last 7 days" = 2026-08-29 through 2026-09-05 inclusive; any deal whose most recent of last_email/last_call/last_meeting is 2026-08-28 or older is stale (8+ days).
- Some engagement rows carry future-dated last_meeting values (e.g. 2026-09-15); per your instruction I used the fields as given, so those deals count as recently contacted. Flagging in case you want future-dated meetings treated differently.
- Verified by script against the two CSVs; working files at /tmp/stale_deals/.
activity-mix-vs-outcomePASS 1.00
Plain-English question: Per rep, what did the last 30 days of emails, calls, and meetings look like next to the deals that advanced?
Prompt — what the model saw
The snapshot date is 2026-09-05. For each sales rep (owner), compute the last-30-day totals of emails, calls, and meetings from the engagements table, and the number of deals that entered DS2 in the last 30 days (t_ds2 within 30 days of the snapshot). Report per rep: the activity mix as percentages (emails/calls/meetings share of that rep's total activities), total activities per DS2 entry, and rank the reps by that efficiency ratio (lowest activities per DS2 entry = most efficient). Name the most efficient rep and the highest-volume rep, and note if they differ. Do not attribute cause.
cost $0.4126 · 350s · in 37,392 / out 18,989 tok
✓ alex_ds2_30d✓ justin_volume_high✓ efficient_named✓ mix_present✓ no_causal_claim

✓ Correct Answer (ground truth)

Plain-English answer: Alex Franklin is the standout signal with 18 deals entering DS2 in 30 days (next closest: Bryce Harmon with 4), alongside per-rep activity totals.
The answer in one line: alex_ds2_30d = 18; bryce_ds2_30d = 4; note = legacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
alex_ds2_30d18
bryce_ds2_30d4
notelegacy alias justin_ds2=4 actually held Bryce Harmon count; Alex Franklin entered 18 deals into DS2 in the 30d window
graded checksalex_ds2_30d · justin_volume_high · efficient_named · mix_present · no_causal_claim

Model's actual answer

show the model's full answer
Snapshot 2026-09-05. Window: t_ds2 in [2026-08-06 .. 2026-09-05] (30 days back, inclusive). Elena Sinclair (701163055) has 1 open deal (Deal-57FF13, DS1, no t_ds2) and no engagement row → 0 activities, 0 DS2 entries; excluded from ranking (no ratio computable).

PER-REP RESULTS (last 30 days)

1. Alex Franklin (84342457) — MOST EFFICIENT
   Emails 307, Calls 36, Meetings 41 → total 384
   Mix: 307/384 = 79.9% emails, 36/384 = 9.4% calls, 41/384 = 10.7% meetings
   DS2 entries: 18 → Deal-403845, Deal-1FC049, Deal-3EED2C, Deal-7FA0C3, Deal-E531A6, Deal-5296C9, Deal-36C33F, Deal-EE195F, Deal-F436DA, Deal-317E6F, Deal-D1E6C2, Deal-D9A72E, Deal-CA5E44, Deal-4F775F, Deal-898FC5, Deal-46988D, Deal-E73427, Deal-92D97D
   Efficiency: 384/18 = 21.3 activities per DS2 entry

2. Bryce Harmon (119337721) — HIGHEST VOLUME per rep? No — see note below
   Emails 162, Calls 0, Meetings 43 → total 205
   Mix: 162/205 = 79.0% emails, 0.0% calls, 43/205 = 21.0% meetings
   DS2 entries: 4 → Deal-25F752, Deal-D73B89, Deal-CA7DC0, Deal-1CCE5C
   Efficiency: 205/4 = 51.3 activities per DS2 entry

3. Cole Ingram (83155924)
   Emails 96, Calls 14, Meetings 1 → total 111
   Mix: 96/111 = 86.5% emails, 14/111 = 12.6% calls, 1/111 = 0.9% meetings
   DS2 entries: 2 → Deal-42326B, Deal-1BA595
   Efficiency: 111/2 = 55.5 activities per DS2 entry

4. Farid Osman (716654662)
   Emails 38, Calls 0, Meetings 34 → total 72
   Mix: 38/72 = 52.8% emails, 0.0% calls, 34/72 = 47.2% meetings
   DS2 entries: 1 → Deal-499BF6
   Efficiency: 72/1 = 72.0 activities per DS2 entry

5. Dana Mercer (83155923)
   Emails 84, Calls 18, Meetings 11 → total 113
   Mix: 84/113 = 74.3% emails, 18/113 = 15.9% calls, 11/113 = 9.7% meetings
   DS2 entries: 1 → Deal-57887A
   Efficiency: 113/1 = 113.0 activities per DS2 entry

EFFICIENCY RANK (lowest activities per DS2 entry first)
1. Alex Franklin — 21.3
2. Bryce Harmon — 51.3
3. Cole Ingram — 55.5
4. Farid Osman — 72.0
5. Dana Mercer — 113.0

HEADLINE
Most efficient rep: Alex Franklin (21.3 activities per DS2 entry).
Highest-volume rep: Alex Franklin (384 total activities, vs. Bryce Harmon 205, Dana Mercer 113, Cole Ingram 111, Farid Osman 72).
They do not differ — the same rep leads on both volume and efficiency.

DATA NOTES
- Archived owners Gavin Porter and Hugo Lindqvist have no deals in the file; excluded.
- Elena Sinclair: 0 activities, 0 DS2 entries (Deal-57FF13 has no engagement row and no t_ds2); no ratio computed.
- No cause attributed; ratios are descriptive only.
ae-qtd-scorecardPASS 1.00
Plain-English question: Build Alex Franklin's QTD scorecard: bookings versus quota, pipeline, and conversion.
Prompt — what the model saw
Build a QTD scorecard for the named AE as of 2026-09-05: bookings vs quota with attainment percent, new vs expansion split, active pipeline by stage with amount, rolling 90-day DS2-to-won rate, win and loss counts with top loss reason, activity volume by type in the last 30 days. Exclude any closed-won deal dated before the quarter. Close with three coaching observations grounded in the numbers.
cost $0.6662 · 466s · in 35,702 / out 34,180 tok
✓ bookings✓ quota✓ attainment✓ new_expansion_split✓ top_loss_reason✓ coaching_obs

✓ Correct Answer (ground truth)

Plain-English answer: 75% attainment ($150,000 on a $200,000 quota), $113,500 new versus $36,500 expansion, pipeline by stage, and the rolling DS2-to-won rate.
The answer in one line: quota = 200000; bookings = 150000; new_bookings = 113500; expansion_bookings = 36500; open_deals_count = 125; attainment_pct = 75; owner = Alex Franklin (+12 more below)
quota200,000
attainment_pct75
bookings150,000
expansion_bookings36,500
expansion_count3
losses_qtd27
new_bookings113,500
new_count5
open_deals_count125
wins_qtd8
activity_30d_totalsemails_30d=807; calls_30d=112; meetings_30d=128; notes_30d=50
allowed_new18 items — e.g. Deal-A1C3E5, Deal-B7D2F4, Deal-C9E1A6, Deal-D4B8C2
ds2_to_wonwindow_start=2026-06-07; won_in_window=8; closed_in_window=35; rate=0.2286
ownerAlex Franklin
pipeline_by_stage{"DS1": {"count": 20, "amount": 284621}, "DS2": {"count": 28, "amount": 353760}, "DS3": {"count": 67, "amount": 552705}, "DS4": {"count": 5, "amount": 23574}, "DS5": {"count": 5, "amount": 45730}}
prior_quarter_excludeddeal_alias=Deal-B3E6F1; close_date=2026-06-20; amount=24000
quarter2026-Q3
snapshot_date2026-09-05
top_loss_reasonreason=Lost- Timing (1 year or more); count=13
graded checksbookings · quota · attainment · new_expansion_split · top_loss_reason · coaching_obs

Model's actual answer

show the model's full answer
AE SCORECARD — Alex Franklin | 2026-Q3 QTD as of 2026-09-05
Quarter elapsed: 67 of 92 days (72.8%)

BOOKINGS VS QUOTA
Bookings QTD: $150,000 (8 closed-won deals)
  = 40,000 + 20,000 + 35,000 + 21,000 + 12,000 + 11,000 + 6,500 + 4,500
Quota (2026-Q3, snapshot 2026-09-05): $200,000
Attainment: 150,000 / 200,000 = 75.0%
Remaining: $50,000 with 25 days left
Excluded pre-quarter: Deal-B3E6F1 ($24,000, closed 2026-06-20) — before 2026-07-01.

NEW VS EXPANSION
New:       5 deals, $113,500 (75.7%) — Deal-A1C3E5, Deal-B7D2F4, Deal-C9E1A6, Deal-D4B8C2, Deal-E6F3A9
Expansion: 3 deals, $36,500 (24.3%) — Deal-F2C7D8, Deal-A8B4D6, Deal-C5D9E2

ACTIVE PIPELINE BY STAGE (125 open deals, $1,260,390 total)
DS1: 20 deals, $284,621
DS2: 28 deals, $353,760
DS3: 67 deals, $552,705
DS4:  5 deals, $23,574
DS5:  5 deals, $45,730
Closing in September (per close_date): $44,367 across 6 deals.

ROLLING 90-DAY DS2-TO-WON (window 2026-05-24 to 2026-08-22, 14-day buffer per standard)
Cohort: 92 deals entered DS2 | Won: 8 | Lost: 27 | Still open: 57
Rate: 8 / 92 = 8.7% (open deals count against the rate)

WIN / LOSS QTD
Won: 8 deals, $150,000 | Lost: 27 deals, $329,272
Win rate by count: 8/35 = 22.9% | by $: 150,000/479,272 = 31.3%
Top loss reason: "Lost- Timing (1 year or more)" — 13 of 27 losses (48%), $184,681 of $329,272 lost (56%)
  Then: MIA 5 ($45,831), Competitor 5 ($49,020), Lost DM 2 ($17,940), Feature Request 1 ($21,000), Does not fit ICP 1 ($10,800)

ACTIVITY, LAST 30 DAYS (all 161 deals in file)
Emails: 807 | Meetings: 128 | Calls: 112 | Notes: 50 | Total: 1,097
On the 8 QTD wins: 3.9 calls + 2.9 meetings per deal. On the 27 QTD losses: 0.9 calls + 0.5 meetings per deal; 8 of 27 losses had zero calls and zero meetings.

COACHING OBSERVATIONS
1. Live engagement separates wins from losses. Won deals averaged 6.8 calls+meetings in 30 days; lost deals averaged 1.4, and 8 of 27 losses (Deal-5885B9, Deal-DC77FE, Deal-55867E, Deal-BF2A98, Deal-583ADB, Deal-E6E80A, Deal-988493, Deal-831B7B) had none at all. The 13 timing losses ($184,681) look like deals that went quiet rather than deals that were beaten — enforce a minimum call/meeting cadence before accepting "timing" as a reason code.
2. The pipeline is big but stale at the top. $1.26M open sounds like 25x the remaining $50K gap, but only $69,304 sits in DS4/DS5, just $44,367 closes in September, and 26 DS2/DS3 deals worth $254,426 entered DS2 before 2026-06-01 (97+ days in stage) — including Deal-F9A08A (entered 2025-11-24) and Deal-901332 (2026-01-28). Clearing or re-qualifying that aged cohort matters more than adding new DS1s; DS2-to-won is running at 8.7% on the rolling 90-day cohort with 57 of 92 still undecided.
3. September is thin — the quarter gets decided in the next two weeks. Only six deals close before quarter-end: Deal-403845 ($9,000), Deal-547B2B ($11,200), Deal-A2B47C ($6,360), Deal-1FC049 ($1,920), Deal-F9A08A ($2,484), Deal-C6D97A ($3,240). Winning all six yields $44,367 against the $50,000 gap — still ~$5,600 short. Hitting quota requires pulling an October close (e.g. Deal-D348E1 $13,770 or Deal-C61CF7 $5,400) forward, or the quarter lands just under 100% even in the best case.
multithreading-gapPASS 1.00
Plain-English question: Which open deals are single-threaded or under-threaded?
Prompt — what the model saw
Find every open deal that is single-threaded (fewer than 2 active contacts) or under-threaded (fewer than 3, or all contacts in one persona). Active means engaged in the last 60 days and not marked former. Personas: economic buyer, champion, HR admin, IT security, finance. For each flagged deal: amount, stage, active contact count, personas present, personas missing, the most valuable persona to add given the stage, and any on-file unengaged contact at that company who fits; otherwise say none on file.
cost $0.0855 · 75s · in 3,409 / out 4,688 tok
✓ flagged_recall✓ persona_language✓ single_thread_term✓ on_file_fit

✓ Correct Answer (ground truth)

Plain-English answer: Deals with fewer than 2 active contacts (single-threaded) or fewer than 3, or all in one persona (under-threaded), each flagged with the missing persona.
The answer in one line: active_window_start = 2026-07-07
active_window_start2026-07-07
allowed_new59 items — e.g. C-077A0E, C-0D15DF, C-1BDFE0, C-2AE3AA
best_persona_ruleDS2/DS3: economic buyer > IT security > finance > HR admin > champion; DS4/DS5: economic buyer > finance > IT security > HR admin > champion; first persona not present among active contacts
flag_ruleflag if active_contacts < 2 (single-threaded) or active_contacts < 3 or all active contacts share one persona (under-threaded); active = engaged within 60 days and not former
flaggeddeal_alias=Deal-36C33F; amount=15000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-50D386; amount=36000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5408B0; amount=14850; stage=DS4; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-5BFE3B; amount=18000; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=economic buyer | deal_alias=Deal-885F45; amount=9300; stage=DS2; active_count=2; reason=under-threaded; best_persona_to_add=IT security | deal_alias=Deal-92D97D; amount=60000; stage=DS2; active_count=1; reason=single-threaded; best_persona_to_add=economic buyer … (+5 more)
healthy_not_flaggedDeal-4B0BEB, Deal-84DBA6, Deal-D348E1
persona_universeeconomic buyer, champion, HR admin, IT security, finance
graded checksflagged_recall · persona_language · single_thread_term · on_file_fit

Model's actual answer

show the model's full answer
DATA GAPS (explicit): Neither file contains deal amount or stage. Amount and stage are therefore unavailable for every deal, and "most valuable persona to add given the stage" is computed stage-blind using this priority rule: champion first if absent, then economic buyer, then IT security, then finance, then HR admin.

ACTIVE WINDOW: Today = 2026-09-24. 60 days back = 2026-07-26. Active = last_engaged_date >= 2026-07-26 AND is_former = false. (2026-06-01 and 2026-06-20 fall outside; 2026-07-30 is inside at 56 days but that contact is former.)

FLAGGED DEALS: 9 of 14. (Not flagged: Deal-84DBA6 — 3 active / 3 personas; Deal-4B0BEB — 4 active / 4 personas; Deal-D348E1 — 5 active / 5 personas.)

1) Deal-EC3025 (C-FDD0C7) — SINGLE-THREADED
   Amount: not provided | Stage: not provided
   Active contacts: 1 (CT-047C54 active; CT-F2C1AE excluded, is_former=true)
   Personas present: champion
   Personas missing: economic buyer, HR admin, IT security, finance
   Add first: economic buyer (champion already present)
   On-file fit: CT-6827DB, Chief People Officer, economic buyer

2) Deal-92D97D (C-E23238) — SINGLE-THREADED
   Amount: not provided | Stage: not provided
   Active contacts: 1 (CT-01F5B4 active; CT-A902AE excluded, last engaged 2026-06-01 = 115 days ago)
   Personas present: HR admin
   Personas missing: champion, economic buyer, IT security, finance
   Add first: champion (none active)
   On-file fit: none on file at C-E23238

3) Deal-50D386 (C-EB10E4) — UNDER-THREADED (2 active)
   Amount: not provided | Stage: not provided
   Active contacts: 2 (CT-AA41B2, CT-B9C35B)
   Personas present: champion, HR admin
   Personas missing: economic buyer, IT security, finance
   Add first: economic buyer
   On-file fit: CT-A1C4B3, Chief People Officer, economic buyer

4) Deal-D0D6B5 (C-32918E) — UNDER-THREADED (3 active, all one persona)
   Amount: not provided | Stage: not provided
   Active contacts: 3 (CT-87CED4, CT-DE6D7C, CT-FD70B2)
   Personas present: champion only
   Personas missing: economic buyer, HR admin, IT security, finance
   Add first: economic buyer
   On-file fit: CT-1FA4DB, Chief People Officer, economic buyer

5) Deal-5BFE3B (C-535D36) — UNDER-THREADED (2 active, all one persona)
   Amount: not provided | Stage: not provided
   Active contacts: 2 (CT-57123B, CT-5CE757)
   Personas present: champion only
   Personas missing: economic buyer, HR admin, IT security, finance
   Add first: economic buyer
   On-file fit: none on file at C-535D36

6) Deal-36C33F (C-077A0E) — SINGLE-THREADED
   Amount: not provided | Stage: not provided
   Active contacts: 1 (CT-4FE556 active; CT-405B45 and CT-86B22F excluded, is_former=true)
   Personas present: IT security
   Personas missing: champion, economic buyer, HR admin, finance
   Add first: champion (none active)
   On-file fit: no champion on file; alternate on-file contact CT-1DB73E, Chief People Officer, economic buyer (fits second-priority gap)

7) Deal-885F45 (C-5E8EFB) — UNDER-THREADED (2 active)
   Amount: not provided | Stage: not provided
   Active contacts: 2 (CT-51C81E, CT-D9A0E8)
   Personas present: economic buyer, champion
   Personas missing: HR admin, IT security, finance
   Add first: IT security (champion + EB already present)
   On-file fit: CT-B3F25D, IT Security Lead, IT security

8) Deal-FCBE5B (C-737030) — SINGLE-THREADED
   Amount: not provided | Stage: not provided
   Active contacts: 1 (CT-4A5317)
   Personas present: champion
   Personas missing: economic buyer, HR admin, IT security, finance
   Add first: economic buyer
   On-file fit: none on file at C-737030

9) Deal-5408B0 (C-2AE3AA) — UNDER-THREADED (2 active)
   Amount: not provided | Stage: not provided
   Active contacts: 2 (CT-D33AE4, CT-8742FD)
   Personas present: champion, HR admin
   Personas missing: economic buyer, IT security, finance
   Add first: economic buyer
   On-file fit: CT-07FA76, Chief People Officer, economic buyer

10) Deal-C6D97A (C-5A8FC2) — UNDER-THREADED (3 active, all one persona)
    Amount: not provided | Stage: not provided
    Active contacts: 3 (CT-223DDC, CT-B03555, CT-4E8A2B)
    Personas present: champion only
    Personas missing: economic buyer, HR admin, IT security, finance
    Add first: economic buyer
    On-file fit: none on file at C-5A8FC2

11) Deal-F9A08A (C-0D15DF) — SINGLE-THREADED
    Amount: not provided | Stage: not provided
    Active contacts: 1 (CT-931B10 active; CT-913581 excluded, last engaged 2026-06-20 = 96 days ago)
    Personas present: champion
    Personas missing: economic buyer, HR admin, IT security, finance
    Add first: economic buyer
    On-file fit: CT-697541, Chief People Officer, economic buyer

SUMMARY ARITHMETIC: 14 deals total; 9 flagged (4 single-threaded: EC3025, 92D97D, 36C33F, FCBE5B; 5 under-threaded: 50D386, D0D6B5, 5BFE3B, 885F45, C6D97A); 3 clean; wait — 9 flagged + 3 clean = 12, not 14. Recount: flagged = EC3025, 92D97D, 50D386, D0D6B5, 5BFE3B, 36C33F, 885F45, FCBE5B, 5408B0, C6D97A, F9A08A = 11 flagged; clean = 84DBA6, 4B0BEB, D348E1 = 3; 11 + 3 = 14. Corrected: 11 flagged (4 single-threaded, 7 under-threaded), 3 clean. Dominant gap: economic buyer missing in 10 of 11 flagged deals; 6 flagged deals have an on-file economic buyer ready to engage (EC3025, 50D386, D0D6B5, 36C33F as alternate, 5408B0, F9A08A).
talk-track-from-transcriptsPASS 1.00
Plain-English question: From this rep's last 10 calls: how do they open, handle objections, and how often do they lock a next step?
Prompt — what the model saw
From this rep's last 10 calls: what they lead with in the first five minutes, how they handle the three most common objections, how often a concrete next step is agreed (rate), and every competitor a prospect raised. Quote at most one line per finding and only lines that appear in the transcripts. End with two coaching notes.
cost $0.0627 · 51s · in 4,262 / out 3,003 tok
✓ objection_count✓ next_step_rate✓ competitors✓ coaching✓ quote_used

✓ Correct Answer (ground truth)

Plain-English answer: 8 of 10 calls open with the same customer-win story; objection handling, next-step agreement rate, and every competitor raised, with quotes.
The answer in one line: n_transcripts = 10; objection_count = 3; rep = Alex Franklin (+3 more below)
n_transcripts10
objection_count3
allowed_new13 items — e.g. TT-001, TT-002, TT-003, TT-004
competitor_quotesAwardco=We're also in late talks with Awardco — their rewards catalog looks bigger than yours.; Kudos=How are you different from Kudos? Our CEO used them at her last company.
competitors_prospect_raisedAwardco, Kudos
next_stepagreed=7; total=10; rate=0.7
objectionsobjection=budget; count=4; quote=Honestly, budget is locked until next fiscal year — I can't add a new line item right now.; transcript_id=TT-001 | objection=timing; count=3; quote=This is good, but can we revisit it next quarter? Open enrollment starts in October and we're underwater.; transcript_id=TT-002 | objection=in-house; count=3; quote=We already do recognition with a spreadsheet and quarterly gift cards — why would we change?; transcript_id=TT-004
opener_themetheme=opens with a customer recognition win story (400-person retailer, 18% turnover reduction); count=8; total=10; quote=Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards, and I wanted to walk you through exactly how they did it.; transcript_id=TT-001
repAlex Franklin
rep_raised_competitors_excludedWorkhuman
graded checksobjection_count · next_step_rate · competitors · coaching · quote_used

Model's actual answer

show the model's full answer
CALL REVIEW: Alex Franklin, TT-001 to TT-010 (10 calls)

1) OPENING (first five minutes)
- 8 of 10 calls (80%) open with the same verbatim proof story: "Quick one to kick off: a 400-person retailer we work with cut regretted turnover 18% in two quarters after automating milestone awards..." (TT-001 / Deal-D348E1; also TT-002, TT-003, TT-005, TT-006, TT-007, TT-008, TT-010)
- Exceptions: TT-004 (Deal-403845) opens agenda-led ("security review first, then pricing"); TT-009 (Deal-1E2498) opens pricing-first.
- In TT-005 (Deal-C61CF7) he adds an unprompted competitor comparison at minute 2 ("And unlike Workhuman, our pricing includes the full rewards catalog with no extra margin.") — rep-raised, not prospect-raised.

2) THREE MOST COMMON OBJECTIONS AND HANDLING
a. Budget locked — 4/10 calls (TT-001, TT-003, TT-006, TT-010). Handled with a turnover-savings reframe: "Most teams fund this out of turnover savings — that retailer saved about $210k in avoided backfills, which is how their finance team signed off."
b. Timing / "revisit next quarter" (open enrollment) — 3/10 (TT-002, TT-005, TT-008). Handled with a de-risked pilot: "What if we scope a 90-day pilot with one department so you have internal data before next quarter's planning?"
c. Status quo (spreadsheet + quarterly gift cards) — 3/10 (TT-004, TT-007, TT-009). Handled with an automation/analytics contrast: "Spreadsheets work until they scale — the difference is automation: milestones fire without HR lifting a finger, and you get analytics on who is being recognized."

3) CONCRETE NEXT-STEP RATE
- Agreed in 7 of 10 calls = 70% (TT-001, TT-002, TT-003, TT-005, TT-006, TT-008, TT-009 — each with a scheduled working session: "Yes, Thursday at 2pm works — send the invite and I'll bring our HRIS manager.").
- No next step in 3 of 10: TT-004 (Deal-403845), TT-007 (Deal-EDC141), TT-010 (Deal-84DBA6).
- Conditional view: when he proposes the working session, acceptance is 7/7 = 100%; the 3 misses are calls where he never proposed one.

4) COMPETITORS RAISED BY PROSPECTS
- Awardco — TT-003 / Deal-547B2B: "We're also in late talks with Awardco — their rewards catalog looks bigger than yours."
- Kudos — TT-007 / Deal-EDC141: "How are you different from Kudos? Our CEO used them at her last company."
- Workhuman appears only in TT-005 and was raised by the rep, not the prospect — excluded from this list per the ask.

COACHING NOTES
1. The opener is word-for-word identical in 8/10 calls; rotate proof points by segment or industry so it doesn't read as scripted, and avoid volunteering competitor comparisons (TT-005) before the prospect asks.
2. On committee/no-urgency stalls he concedes with no path forward ("Understood — I'll leave it with you.", "Fair enough.") and those calls close 0/3 on next steps; extend the existing reframe/pilot playbook to these objections (e.g., propose a committee-readiness session), since every time he does ask for a concrete step, he gets it (7/7).
reporting-analytics 5 tests
q3-weighted-forecastPASS 1.00
Plain-English question: Build the Q3 2026 weighted forecast from the open pipeline, and show the math deal by deal.
Prompt — what the model saw
Below are the open deals in the default sales pipeline with a close date between 2026-07-01 and 2026-10-15, pulled from the CRM warehouse on 2026-09-05. Stages: DS1 (discovery) through DS5 (contract). Build the Q3 2026 forecast (quarter = 2026-07-01 to 2026-09-30). Weighting rule: 100 percent of COMMIT plus 35 percent of BEST_CASE. PIPELINE deals count zero. Only deals with a close date inside the quarter count. Report: COMMIT total, BEST_CASE total, weighted forecast, count of deals in each category inside the quarter, the deals you excluded for being outside the quarter (count and total amount), and the top 5 BEST_CASE deals by amount inside the quarter. Show arithmetic. Then, in a final '## Data quality' section of no more than four sentences, name the data quality problems in this extract that would make the forecast unreliable if it were run unattended.
cost $0.1794 · 148s · in 8,407 / out 9,537 tok
✓ commit_total_44729✓ best_case_total_203565✓ weighted_115976.75✓ commit_count_7✓ best_case_count_24✓ excluded_count_32✓ excluded_amount_227575✓ top5_aliases✓ flags_blank_owner✓ flags_missing_why_buys

✓ Correct Answer (ground truth)

Plain-English answer: A stage-weighted Q3 forecast number computed from the open pipeline, with per-stage math that reconciles to the deal data.
The answer in one line: excluded_amount = 227575; best_case_total = 203565; weighted_forecast = 115,976.75; commit_total = 44729; excluded_out_of_quarter = 32; best_case_count = 24; flags = blank owner column + empty why_buys fields must be surfaced (+2 more below)
excluded_amount227,575
best_case_count24
best_case_total203,565
commit_count7
commit_total44,729
excluded_out_of_quarter32
weighted_forecast115,976.75
flagsblank owner column + empty why_buys fields must be surfaced
top5_best_caseDeal-2D7423, Deal-25F752, Deal-E53952, Deal-5EED42, Deal-FA32A0
graded checkscommit_total_44729 · best_case_total_203565 · weighted_115976.75 · commit_count_7 · best_case_count_24 · excluded_count_32 · excluded_amount_227575 · top5_aliases · flags_blank_owner · flags_missing_why_buys

Model's actual answer

show the model's full answer
Q3 2026 FORECAST (2026-07-01 to 2026-09-30)

In-quarter deals: 54 of 86 (COMMIT 7, BEST_CASE 24, PIPELINE 23)

COMMIT total: $44,729
  11,200 (Deal-547B2B) + 9,000 (Deal-B7EBD1) + 9,000 (Deal-403845) + 6,360 (Deal-A2B47C) + 5,400 (Deal-2465CE) + 2,520 (Deal-A5E80A) + 1,249 (Deal-499BF6) = 44,729

BEST_CASE total: $203,565 (24 deals)

Weighted forecast = 100% x COMMIT + 35% x BEST_CASE
  = 44,729 + (0.35 x 203,565)
  = 44,729 + 71,247.75
  = $115,976.75

Excluded for close date outside the quarter: 32 deals, $227,575 total
  (all after 2026-09-30: 1 COMMIT $13,770, 9 BEST_CASE $28,240, 22 PIPELINE $185,565 — none before 2026-07-01)

Top 5 BEST_CASE deals in-quarter by amount:
  1. Deal-2D7423 — $38,935
  2. Deal-25F752 — $24,000
  3. Deal-E53952 — $19,656
  4. Deal-5EED42 — $16,250
  5. Deal-FA32A0 — $11,116

## Data quality
85 of 86 deals have a blank owner field, so no accountability check or per-rep roll-up is possible. Stage and forecast_category conflict on several in-quarter deals — Deal-A5E80A (DS1/COMMIT), Deal-499BF6 (DS2/COMMIT), Deal-C61CF7 (DS5/BEST_CASE) — meaning category assignment isn't trustworthy. Four deals (Deal-333EBB 2026-08-28, Deal-57FF13 2026-09-02, Deal-31AD2C 2026-09-04, Deal-7A2454 2026-09-04) carry close dates already in the past relative to the 2026-09-05 pull date, indicating stale close dates. why_buys_chars is 0 on 71 of 86 deals (including all 7 in-quarter COMMIT deals), and one amount is fractional (Deal-42326B = 2480.4), suggesting incomplete or unvalidated field population.
aha-moment-2x2PASS 1.00
Plain-English question: Do companies that hit 5+ givers and one redemption in month one actually retain better at 24 months?
Prompt — what the model saw
Test the activation hypothesis on the company cohort below: companies with 5+ unique givers (m1_users >= 5) AND 1+ successful redemption (m1_redemptions >= 1) in their first calendar month retain better at 24 months than those with one signal or neither. Every company in this extract is already 25+ months old, so current_status = 'active' means retained at 24 months. Report the full 2x2 (cohort size and 24-month retention rate for: both signals, givers-only, redemption-only, neither), state how many companies were excluded from the denominator and why (if none, say so), name the single signal with the largest retention lift, and state what this does and does not prove.
cost $0.3472 · 213s · in 41,922 / out 14,107 tok
✓ both_66✓ givers_47✓ redemp_31✓ neither_40✓ excluded_0✓ largest_lift_givers

✓ Correct Answer (ground truth)

Plain-English answer: The four cohort cells are graded against computed 24-month retention; the numbers come from the scorer's cohort table, not opinion.
The answer in one line: both_5plus_givers_and_redemption = 66.0% (31 of 47)
both_5plus_givers_and_redemption66.0% (31 of 47)
exclusionsnone — claim bounded as correlation, not causation
givers_only46.9% (23 of 49)
largest_single_signal_liftunique givers
neither40.0% (38 of 95)
redemption_only31.0% (9 of 29)
graded checksboth_66 · givers_47 · redemp_31 · neither_40 · excluded_0 · largest_lift_givers

Model's actual answer

show the model's full answer
Activation hypothesis test — 2x2 on plg_company_cohort_2023.csv
Note: the skill for this data applies here; per its company-level aha definition, m1_users is treated as the "5+ unique givers" signal exactly as the hypothesis specifies (m1_users >= 5) and m1_redemptions >= 1 as the redemption signal. Retained = current_status = 'active'.

FULL 2x2 (n=220, all statuses in denominator)

  Cohort                       n    Retained   24-mo retention
  Both signals (5+ givers + 1+ redemption)   47   31   31/47  = 66.0%
  Givers-only (5+ users, 0 redemptions)      49   23   23/49  = 46.9%
  Redemption-only (<5 users, 1+ redemption)  29    9    9/29  = 31.0%
  Neither                                    95   38   38/95  = 40.0%

  Overall: 101/220 = 45.9% retained.

EXCLUSIONS FROM THE DENOMINATOR: none. All 220 rows were included; no duplicate company_keys. Data caveat, not an exclusion: 3 companies are current_status='non_renewing' (C-0B2078FB [neither], C-0A96134F [redemption-only], C-0BEAF685 [redemption-only]). They are counted as NOT retained in the primary table above, since they are not active at 24 months. Sensitivity check excluding them from the denominator entirely: both 66.0% (31/47), givers-only 46.9% (23/49), redemption-only 33.3% (9/27), neither 40.4% (38/94) — directionally unchanged.

ARITHMETIC / LIFTS (vs neither, primary table)

  Both vs neither:            66.0% - 40.0% = +26.0 pp
  Givers-only vs neither:     46.9% - 40.0% = +6.9 pp
  Redemption-only vs neither: 31.0% - 40.0% = -9.0 pp
  Both vs givers-only:        66.0% - 46.9% = +19.0 pp
  Both vs redemption-only:    66.0% - 31.0% = +34.9 pp

SINGLE SIGNAL WITH THE LARGEST RETENTION LIFT: the givers signal (m1_users >= 5). Givers-only beats neither by +6.9 pp, while redemption-only actually underperforms neither by 9.0 pp. Redemption alone is a negative-selection signal here — companies that redeem without reaching 5 givers retain worst of any cell (31.0%).

WHAT THIS PROVES
- Directionally supports the hypothesis: the both-signals cohort retains at 66.0%, the highest of all four cells, +26.0 pp over neither and +19.0/+34.9 pp over each single-signal cell.
- The two signals appear complementary: breadth of participation (5+ givers) carries positive lift on its own; redemption only pays off when breadth is also present.

WHAT THIS DOES NOT PROVE
- Causation. This is an observational cohort split, not an experiment. Companies that hit both signals may simply be better-fit, larger, or better-onboarded accounts (self-selection); the extract includes no controls for company size, industry, tier, or acquisition channel.
- Statistical significance. Cells are small (n=29 to 95); a +6.9 pp lift on n=49 vs n=95 is well within sampling noise, and no confidence intervals or significance tests were run. The redemption-only penalty (-9.0 pp, n=29) is likewise not established as real.
- That the thresholds are optimal. 5 users and 1 redemption were tested as given; the data says nothing about whether 3 users, 10 users, or 3 redemptions would discriminate better.
- Generalizability. This is one 2023 signup extract (Jan–Jul 2023, n=220); results may not hold for other cohort windows or for the full self-serve population, and the extract's 45.9% overall retention cannot be compared to a company-wide baseline from this file alone.
- Treatment of non_renewing as non-retained is a definitional choice; results are robust to it, but the file gives no cancellation dates, so "retained at exactly month 24" cannot be verified — only status as of extract date.
arr-rollforward-reconciliationPASS 1.00
Plain-English question: Does CRM ARR match billing ARR — and if not, exactly where is the variance?
Prompt — what the model saw
Reconcile ARR between CRM company records and billing subscriptions as of 2026-09-05. Billing ARR is MRR x 12. Report both totals, the variance, and a decomposition into buckets (status mismatch, rounding, missing records, other) that sums exactly to the variance. List each mismatched account with a suggested owner. Business rule: any subscription with a term other than 12 months must have cf_agreement_end_date populated; list violations.
cost $0.0666 · 58s · in 3,151 / out 3,481 tok
✓ billing_total✓ crm_total✓ variance✓ buckets_sum✓ violations_listed

✓ Correct Answer (ground truth)

Plain-English answer: Billing $604,739.28 vs CRM $603,581.76: a $1,157.52 variance decomposed into status-mismatch, rounding, and missing-record buckets.
The answer in one line: n_companies = 40; n_mismatched_accounts = 7; as_of = 2026-09-05 (+4 more below)
n_companies40
n_mismatched_accounts7
as_of2026-09-05
billing_arr_total604739.28
bucket_sum_equals_variance1
buckets{"status_mismatch": {"amount": "-13158.48", "accounts": [{"company_alias": "C-0C8323BF", "billing_arr": "0.00", "hubspot_arr": "4905.24", "diff": "-4905.24", "suggested_owner": "Dana Mercer"}, {"compa
crm_arr_total603581.76
ruleterm_months != 12 requires cf_agreement_end_date
variance1157.52
variance_definitionbilling_arr_total - crm_arr_total (billing ARR = sum of MRR x 12 over status=active subscriptions)
violationssubscription_id=SUB-0002; company_alias=C-1794A52C; term_months=24 | subscription_id=SUB-0019; company_alias=C-22170CA1; term_months=36
graded checksbilling_total · crm_total · variance · buckets_sum · violations_listed

Model's actual answer

show the model's full answer
ARR RECONCILIATION — as of 2026-09-05

Totals
  Billing ARR (active subs only, MRR x 12, 36 subscriptions): $604,739.28
  CRM ARR (HubSpot company records, 40 companies):            $603,581.76
  Variance (Billing - CRM):                                   $1,157.52

Arithmetic notes: SUB-000E and SUB-000F are status=cancelled, so excluded from billing ARR. CRM total includes all 40 rows as given.

Variance decomposition (sums exactly to $1,157.52)

  1. Status mismatch: -$13,158.48
     Cancelled subscriptions still carrying CRM ARR:
       C-0C8323BF (SUB-000E): CRM $4,905.24, billing $0 -> -4,905.24
       C-0DC4FB8C (SUB-000F): CRM $8,253.24, billing $0 -> -8,253.24

  2. Missing records: +$11,952.00
       C-21629AA4 (SUB-0004): billing $28,449.24 (2,370.77 x 12), no CRM company record -> +28,449.24
       C-0D5BBE3A: CRM $16,497.24, no billing subscription -> -16,497.24

  3. Rounding: -$36.00
       C-0D66DF9E (SUB-0005): billing 1,932.00 x 12 = $23,184.00 vs CRM $23,200.00 -> -16.00
       C-14D70CE0 (SUB-0008): billing 1,515.00 x 12 = $18,180.00 vs CRM $18,200.00 -> -20.00

  4. Other: +$2,400.00
       C-0F7269D7 (SUB-0006): billing 2,233.00 x 12 = $26,796.00 vs CRM $24,396.00 -> +2,400.00
       Note: $2,400 = $200 x 12, consistent with a $200/mo MRR discrepancy (data entry or unrecorded plan change), not rounding.

  Check: -13,158.48 + 11,952.00 - 36.00 + 2,400.00 = 1,157.52 ✓

Mismatched accounts (all 32 other accounts tie out to the cent)

  Account       Billing ARR  CRM ARR    Diff        Bucket            Suggested owner
  C-21629AA4    28,449.24    (none)     +28,449.24  Missing (CRM)     RevOps / CRM admin — create company record
  C-0D5BBE3A    (none)       16,497.24  -16,497.24  Missing (billing) Billing ops — locate/restore subscription
  C-0C8323BF    0.00         4,905.24   -4,905.24   Status mismatch   CS owner — zero out CRM ARR on churn
  C-0DC4FB8C    0.00         8,253.24   -8,253.24   Status mismatch   CS owner — zero out CRM ARR on churn
  C-0F7269D7    26,796.00    24,396.00  +2,400.00   Other             AE of record + billing ops — verify MRR
  C-0D66DF9E    23,184.00    23,200.00  -16.00      Rounding          RevOps — align ARR rounding convention
  C-14D70CE0    18,180.00    18,200.00  -20.00      Rounding          RevOps — align ARR rounding convention

Owner names are not present in the data, so owners above are suggested by role only.

Business rule violations (term ≠ 12 months requires cf_agreement_end_date)

  SUB-0002 (C-1794A52C): term 24 months, cf_agreement_end_date empty — VIOLATION
  SUB-0019 (C-22170CA1): term 36 months, cf_agreement_end_date empty — VIOLATION

  Compliant non-12-month subs: SUB-000C (24 mo, 2027-11-30), SUB-001A (36 mo, 2027-11-30).
kvm-month-over-monthPASS 1.00
Plain-English question: How did the four core KVMs move in August versus July?
Prompt — what the model saw
Report the four core KVMs for 2026-08 against 2026-07: giving rate, redemptions per user, 1:1 meetings engagement, pulse check engagement. For each: value, prior value, absolute and relative change, direction. Name the KVM with the largest relative move and the segment (plan_tier or size_band) driving it if the data supports that; otherwise say the data does not. Under 200 words plus one table.
cost $0.1444 · 85s · in 20,449 / out 5,538 tok
✓ four_kvms✓ biggest_mover_named✓ direction_correct✓ segment_named✓ word_cap

✓ Correct Answer (ground truth)

Plain-English answer: Giving rate 60.27% (up 0.07%), redemptions per user 1.7302 (up 0.01%), with 1:1 meetings and pulse engagement each up fractionally — direction reported for all four.
The answer in one line: n_company_rows_per_month = 30; aggregation = simple mean of company-level values per month
n_company_rows_per_month30
aggregationsimple mean of company-level values per month
biggest_moverpulse_engagement
driving_segmentsize_band=enterprise; 2026-07=0.55; 2026-08=0.2743; rel_change_pct=-50.13
kvms{"giving_rate": {"2026-07": 0.6023, "2026-08": 0.6027, "abs_change": 0.0004, "rel_change_pct": 0.07, "direction": "up"}, "redemptions_per_user": {"2026-07": 1.73, "2026-08": 1.7302, "abs_change": 0.00
segment_detail{"smb": {"2026-07": 0.6588, "2026-08": 0.6573, "rel_change_pct": -0.23}, "mid_market": {"2026-07": 0.593, "2026-08": 0.5942, "rel_change_pct": 0.2}, "enterprise": {"2026-07": 0.55, "2026-08": 0.2743,
graded checksfour_kvms · biggest_mover_named · direction_correct · segment_named · word_cap

Model's actual answer

show the model's full answer
Method: unweighted mean across the same 30 companies each month (10 smb, 10 mid_market, 10 enterprise; all tier_three).

| KVM | 2026-07 | 2026-08 | Abs chg | Rel chg | Direction |
|---|---|---|---|---|---|
| Giving rate | 0.6023 | 0.6027 | +0.0004 | +0.07% | Up (flat) |
| Redemptions/user | 1.7300 | 1.7302 | +0.0002 | +0.01% | Up (flat) |
| 1:1 meetings engagement | 0.4469 | 0.4472 | +0.0003 | +0.06% | Up (flat) |
| Pulse check engagement | 0.6006 | 0.5086 | -0.0920 | -15.32% | Down |

Arithmetic (pulse): Jul = (6.5879 smb + 5.9299 mm + 5.4998 ent)/30 = 18.0176/30 = 0.6006. Aug = (6.5731 + 5.9424 + 2.7428)/30 = 15.2583/30 = 0.5086. Change = -0.0920; -0.0920/0.6006 = -15.32%.

Largest relative move: pulse check engagement, -15.32%. Driver: the enterprise size_band. Enterprise pulse fell 0.5500 → 0.2743 (-0.2757, -50.1%), with all 10 enterprise accounts dropping sharply (e.g. C-0B2895EF 0.5619 → 0.2864). SMB (0.6588 → 0.6573, -0.22%) and mid_market (0.5930 → 0.5942, +0.21%) were flat. plan_tier cannot explain the move — every record is tier_three.

Caveat: these are unweighted company averages; the data has no user counts, so account-size weighting effects cannot be assessed.
rewards-ytd-sectionPASS 1.00
Plain-English question: Produce the redemption section of the weekly rewards report through August.
Prompt — what the model saw
Produce the redemption section of the weekly rewards report through the last completed month (2026-08): redemption count, spend, unique redeemers, redemptions per redeemer, provider mix as percent of spend, top 5 countries by redemptions. Say which month is the last completed month. Provider shares must sum to 100.
cost $0.5266 · 266s · in 69,564 / out 17,844 tok
✓ count✓ spend✓ unique_redeemers✓ month_named✓ provider_mix✓ top_country

✓ Correct Answer (ground truth)

Plain-English answer: 378 redemptions, $27,846 spend, 235 unique redeemers at 1.61 each, custom rewards at 39% of spend, US leading all countries.
The answer in one line: redemption_count = 378; unique_redeemers = 235; redemptions_per_redeemer = 1.61; total_spend_usd = 27846.00
redemption_count378
redemptions_per_redeemer1.61
unique_redeemers235
provider_mix_pct_of_spendTangoCard=11.6; Tremendous=30.54; custom=39.05; Snappy=18.81
top5_countries_by_countcountry=US; count=244 | country=CA; count=24 | country=AU; count=21 | country=GB; count=17 | country=NL; count=17
total_spend_usd27846.00
graded checkscount · spend · unique_redeemers · month_named · provider_mix · top_country

Model's actual answer

show the model's full answer
REDEMPTIONS — YTD Jan 1 – Aug 31, 2026

Last completed month: 2026-08 (August 2026). The file contains no redemptions after 2026-08-31, so all 378 rows fall inside the reporting window; September is not represented at all (not a partial month in this data).

Volume & spend
- Redemptions: 378
- Spend: $27,846.00 (sum of amount_usd)
- Unique redeemers: 235 (distinct user_key, deduped YTD — not a sum of monthly uniques)
- Redemptions per redeemer: 378 / 235 = 1.61

Provider mix (% of spend)
- custom: $10,873.00 → 10,873 / 27,846 = 39.0% (37 redemptions)
- Tremendous: $8,505.00 → 8,505 / 27,846 = 30.5% (192 redemptions)
- Snappy: $5,238.00 → 5,238 / 27,846 = 18.8% (59 redemptions)
- TangoCard: $3,230.00 → 3,230 / 27,846 = 11.6% (90 redemptions)
- Shares: 39.0 + 30.5 + 18.8 + 11.6 = 99.9 at 1 decimal; exact shares are 39.05 / 30.54 / 18.81 / 11.60, which sum to 100.0. Rounding one share to force 100.0 at 1dp: custom 39.1 / Tremendous 30.5 / Snappy 18.8 / TangoCard 11.6 = 100.0.

Top 5 countries by redemptions
1. US — 244 ($18,547.00)
2. CA — 24 ($2,286.00)
3. AU — 21 ($1,606.00)
4. GB — 17 ($944.00)
5. NL — 17 ($1,122.00)

Note: GB and NL tie at 17 redemptions; both shown. Next is SG at 12.
customer-success 5 tests
churn-save-eligibilityPASS 1.00
Plain-English question: Which at-risk accounts actually qualify for a churn-save offer, and which play fits each?
Prompt — what the model saw
Which at-risk accounts qualify for a churn-save offer under the documented eligibility rules, what amount is at stake per account and in total, and which play fits each (usage revival, executive touch, commercial concession)? Cite the signal that justifies each play. List accounts that look at risk but do not qualify and why.
cost $0.1683 · 109s · in 15,354 / out 6,848 tok
✓ eligible_set✓ total_at_stake✓ plays_cited✓ noneligible_named✓ rules_applied

✓ Correct Answer (ground truth)

Plain-English answer: Only accounts meeting all three rules qualify — e.g. C-0F6C0F34 at $49,707 (executive touch) and C-0B827671 at $25,365 (usage revival) — with amount and play per account.
The answer in one line: n_accounts = 30; snapshot_date = 2026-09-05
n_accounts30
eligibleaccount_alias=C-0F6C0F34; amount_at_stake=49707.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0B827671; amount_at_stake=25365.00; play=usage revival; justifying_signal=usage_trend_3m=declining over the last 3 months | account_alias=C-0B360C78; amount_at_stake=35748.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 75% - risk is commercial, not adoption | account_alias=C-0B0F1BAB; amount_at_stake=5494.00; play=executive touch; justifying_signal=champion_active is false - no executive sponsor engaged | account_alias=C-0CA21961; amount_at_stake=16829.00; play=usage revival; justifying_signal=seat utilization 26% is below 50% | account_alias=C-0E9C27D1; amount_at_stake=41235.00; play=commercial concession; justifying_signal=usage stable/growing with seat utilization 85% - risk is commercial, not adoption … (+2 more)
non_eligible_at_riskaccount_alias=C-0BC71BDD; health_score=55 | account_alias=C-0BA71F12; health_score=52 | account_alias=C-0F6694C3; health_score=43 | account_alias=C-0BE96399; health_score=54 | account_alias=C-0F876796; health_score=47 | account_alias=C-0FCCD2DF; health_score=43 … (+1 more)
ruleshealth_score < 60, churn_save_eligible_amount > 0, renewal within 120 days of snapshot
snapshot_date2026-09-05
total_amount_at_stake224601.00
graded checkseligible_set · total_at_stake · plays_cited · noneligible_named · rules_applied

Model's actual answer

show the model's full answer
ELIGIBILITY — all three rules must pass (snapshot 2026-09-05; R3 cutoff = 2026-09-05 + 120 days = 2027-01-03).

QUALIFY FOR CHURN-SAVE OFFER (8 accounts, all pass R1 health<60, R2 eligible amount>0, R3 renewal within 120 days)

Account | At-stake (eligible) | Renewal (days out) | Play | Justifying signal
C-0F6C0F34 | $49,707.00 | 2026-10-03 (28d) | Executive touch | champion_active=false, usage growing (78% of seats) — product is landing but no executive sponsor with renewal 28 days out
C-0E9C27D1 | $41,235.00 | 2026-09-24 (19d) | Executive touch + commercial concession | health 39 (lowest-tier), renewal in 19 days, but champion active and 85% seat utilization — relationships and usage are intact; the gap is time and likely commercial, so pair exec alignment with a concession to close fast
C-0B360C78 | $35,748.00 | 2026-10-28 (53d) | Commercial concession | usage growing, champion active, 75% utilization — no usage or sponsor gap; health 57 with strong fundamentals points to a commercial blocker
C-0CEF69FD | $32,621.00 | 2026-11-21 (77d) | Executive touch | champion_active=false with usage growing (71% utilization) — adoption without an executive sponsor
C-0B827671 | $25,365.00 | 2026-11-14 (70d) | Usage revival | usage_trend=declining, 56% seat utilization (113/202); champion still active to sponsor the re-engagement
C-0D3278C7 | $17,602.00 | 2026-11-12 (68d) | Usage revival | usage_trend=declining, 33% utilization (126/380 seats) — biggest seat waste in the qualified set
C-0CA21961 | $16,829.00 | 2026-12-28 (114d) | Usage revival | 26% utilization (84/325 seats), trend flat — deepest under-deployment; champion active to drive it
C-0B0F1BAB | $5,494.00 | 2026-09-23 (18d) | Executive touch | health 38, champion_active=false, renewal in 18 days — no time for a usage motion; needs immediate sponsor-level save

TOTAL AT STAKE (eligible amount): $224,601.00
Arithmetic: 49,707 + 41,235 + 35,748 + 32,621 + 25,365 + 17,602 + 16,829 + 5,494 = 224,601
Combined ARR of the 8 qualified accounts: $454,380.00 (86,741 + 75,093 + 60,427 + 79,324 + 72,088 + 33,815 + 31,501 + 15,391)

AT RISK (health < 60) BUT DO NOT QUALIFY (7 accounts)

C-0BE96399 (h=54) — fails R2: eligible amount $0.00 despite declining usage, 28% utilization, renewal 54 days out
C-10A56B0F (h=54) — fails R2: eligible amount $0.00 despite declining usage, renewal 98 days out
C-0BC71BDD (h=55) — fails R2: eligible amount $0.00; 30% utilization, champion inactive, renewal 52 days out
C-0F876796 (h=47) — fails R3: renewal 2027-02-06 is 154 days out (>120). Eligible amount $19,958.00 exists; declining usage, 23% utilization, champion inactive — strongest watchlist candidate
C-0BA71F12 (h=52) — fails R3: renewal 2027-04-11 is 218 days out; eligible amount $6,824.00, declining usage, 23% utilization
C-0F6694C3 (h=43) — fails R2 and R3: eligible amount $0.00, renewal 197 days out
C-0FCCD2DF (h=43) — fails R2 and R3: eligible amount $0.00, renewal 230 days out

Note: the play taxonomy (usage revival / executive touch / commercial concession) is not defined in the provided files; assignments above are inferred from the signals in churnzero_accounts.csv and each signal is cited per account.
expansion-kitPASS 1.00
Plain-English question: Build the expansion kit for the customer who just replied to our outreach.
Prompt — what the model saw
A customer replied to outreach (reply included). Build the expansion kit: seat coverage (licensed over headcount), two lines on usage health, headroom in seats and ARR at the current per-seat rate, who replied and whether they can buy, the right buyer from our contacts, and a reply email under 150 words that moves toward the right person without being pushy and cites one real usage fact.
cost $0.0358 · 38s · in 2,041 / out 1,651 tok
✓ coverage_pct✓ headroom_arr✓ right_buyer✓ replier_not_buyer✓ usage_fact✓ email_under_150

✓ Correct Answer (ground truth)

Plain-English answer: 37.5% seat coverage, 250 headroom seats worth $15,000 ARR at $60 a seat, usage up 43.2% in six months — and the replier cannot buy; Dana R. (VP People) is the right buyer.
The answer in one line: headroom_seats = 250; coverage_pct = 37.5; company_alias = C-0DDFC9A7 (+3 more below)
headroom_seats250
coverage_pct37.5
company_aliasC-0DDFC9A7
headroom_arr15000.00
per_seat_annual_rate60.00
replierMaria S.
replier_can_buy0
right_buyername=Dana R.; title=VP People
usage_factsactive_users_2026_08=126; active_users_2026_03=88; growth_pct_6m=43.2; licensed_seat_utilization_pct=84
usage_health_linesMonthly active users grew from 88 (2026-03) to 126 (2026-08), +43.2% over 6 months., Last month 126 of 150 licensed seats were active (84.0% utilization), approaching the seat cap.
graded checkscoverage_pct · headroom_arr · right_buyer · replier_not_buyer · usage_fact · email_under_150

Model's actual answer

show the model's full answer
EXPANSION KIT — C-0DDFC9A7

Seat coverage
150 licensed / 400 headcount = 37.5% coverage. 62.5% of the company is outside the program.

Usage health (two lines)
1. MAU has grown every month on record: 88 (Mar) → 95 → 102 → 110 → 118 → 126 (Aug), a +43.2% increase over the period ((126−88)/88).
2. August active users (126) now equal 84% of licensed seats (126/150); at the recent pace of ~7–8 new actives/month, seats fully saturate in roughly 3 months — usage is pressing against the license cap.

Headroom
- Seats: 400 headcount − 150 licensed = 250 seats of headroom.
- Per-seat rate: $9,000.00 ARR / 150 seats = $60.00/seat/yr.
- ARR headroom: 250 seats × $60.00 = $15,000.00 expansion ARR; $24,000.00 potential total ARR at full headcount coverage.

Who replied / can they buy
Maria S., People Operations Coordinator (last engaged 2026-09-02, the reply date). She explicitly says she is NOT the buyer: "Budget and seat expansion sit with Dana R." She offered to make an introduction and noted Dana "has been asking about our usage numbers lately."

Right buyer
Dana R., VP People — named by Maria as the budget/seat-expansion owner. Caveat: Dana's last engagement was 2026-05-18, ~4 months stale, so the warm intro via Maria is the right path rather than a cold direct touch. Sam K. (Office Manager, last engaged 2025-11-03) is not relevant.

Reply email (107 words)

Hi Maria,

Thanks — glad to hear the feed stays busy. The numbers back it up: your monthly active users grew from 88 in March to 126 in August, which is about 84% of your 150 licensed seats.

An introduction to Dana would be great, whenever the timing works for you — no rush. Since you mentioned she's been asking about usage, I'll come prepared with a short summary she can use internally.

If it's easier, I'm happy to send that summary to you first so you can share it directly — whichever you prefer.

Thanks again,
Cole
onboarding-statusPASS 1.00
Plain-English question: What is done and what is missing in this customer's onboarding before the mid-onboarding call?
Prompt — what the model saw
Prep the CSM for the mid-onboarding call: what is complete (integrations connected, allowance set, admins added, first recognition given, first redemption), what is not, early engagement signals, and the three things to cover on the call. Mark nothing complete without a data field that shows it.
cost $0.0265 · 26s · in 2,043 / out 1,029 tok
✓ complete_items✓ incomplete_items✓ three_items✓ grounded

✓ Correct Answer (ground truth)

Plain-English answer: 25 days live: Slack connected, allowance set, admins added — but the HRIS integration is not connected, with early engagement signals and three call priorities.
The answer in one line: days_live = 25; company_alias = C-0D284E42
days_live25
call_itemsConnect the HRIS integration (integration_hris is empty) to automate user provisioning and offboarding., Get to first redemption (first_redemption_at is empty): walk admins through the rewards catalog and confirm allowance funding., Build the giving habit: daily active givers averaged 4.3 in week 1 vs 13.0 last week - agree on manager nudges and a recognition cadence.
checklist{"slack_integration_connected": {"complete": true, "field": "integration_slack", "value": "2026-08-12"}, "hris_integration_connected": {"complete": false, "field": "integration_hris", "value": ""}, "a
company_aliasC-0D284E42
engagement_signalsavg_active_givers_week1=4.3; avg_active_givers_last_week=13; latest_active_givers=15; days_with_activity=25
graded checkscomplete_items · incomplete_items · three_items · grounded

Model's actual answer

show the model's full answer
Mid-onboarding prep — C-0D284E42 (signed up 2026-08-11; usage data through 2026-09-04, 24 days in)

COMPLETE (data-backed)
- Slack integration: connected 2026-08-12 (1 day post-signup)
- Allowance set: 2026-08-13 (2 days post-signup)
- Admins added: 2
- First recognition given: 2026-08-15 14:22 (4 days post-signup)

NOT COMPLETE (no data field shows it)
- HRIS integration: integration_hris is empty — not connected
- First redemption: first_redemption_at is empty — no recipient has redeemed after 24 days

EARLY ENGAGEMENT SIGNALS
- Active givers grew from 3 (2026-08-11) to 15 (2026-09-03 and 2026-09-04, the peak)
- First full week avg (Aug 11–17): (3+3+4+4+5+4+7)/7 = 30/7 = 4.3 givers/day
- Last full week avg (Aug 29–Sep 4): (11+13+11+13+13+15+15)/7 = 91/7 = 13.0 givers/day
- That's a 3.0x increase in daily active givers over ~3 weeks, with no down-week trend — adoption is compounding on the giving side
- Caveat: data shows givers only. No recipient-activity or redemption-side data was provided, so I can't confirm the receiving side of the loop is engaging

THREE THINGS TO COVER ON THE CALL
1. HRIS integration — still not connected (empty field). This is the biggest onboarding gap: without it, user provisioning/sync is manual. Get an owner and a target date.
2. Zero redemptions — 24 days of rising giving (peak 15 givers) but first_redemption_at is empty. Redemption is the value moment for recipients. Check: is the rewards catalog configured and visible, is the allowance amount meaningful against catalog prices, do recipients know they have a balance?
3. Sustain and widen the giving momentum — engagement is 3x and climbing, but it rests on 2 admins and giving-side activity only. Confirm whether they want more admins/manager champions and a comms plan to pull recipients (and non-givers) in before the giving cohort plateaus.
renewal-risk-conflicting-datesPASS 1.00
Plain-English question: Two systems disagree on renewal dates — which do you trust per account, and who is actually renewing inside 90 days?
Prompt — what the model saw
Produce a 90-day renewal risk brief. Two systems hold renewal dates and disagree on some accounts; decide which to trust per account and say why (multi-year contracts are known to be wrong in ChurnZero). For every renewal: company, CSM, ARR, date used, seat utilization, 3-month usage trend, risk rating with one sentence of evidence. Flag every disagreement. Close with total ARR renewing and ARR at risk.
cost $0.1456 · 125s · in 6,475 / out 8,081 tok
✓ total_renewing✓ arr_at_risk✓ disagreements_flagged✓ trust_rule

✓ Correct Answer (ground truth)

Plain-English answer: Chargebee wins per account (multi-year contracts are known wrong in ChurnZero), with each trusted date, the reason, and a 90-day-window flag.
The answer in one line: n_accounts = 20; n_disagreements = 5; snapshot_date = 2026-09-05 (+2 more below)
n_accounts20
n_disagreements5
accounts20 items — e.g. account_alias=C-0B144C78; csm=Cole Ingram; arr=30899.00; trusted_renewal_date=2026-11-02; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.4; usage_3m_ratio=1.03; risk=low; evidence=3-month usage ratio 1.03 (last3 avg 103 vs prior3 100), seat utilization 75% | account_alias=C-0B20DB64; csm=Dana Mercer; arr=21770.00; trusted_renewal_date=2026-10-07; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=56.6; usage_3m_ratio=1; risk=medium; evidence=3-month usage ratio 1.00 (last3 avg 295 vs prior3 295), seat utilization 57% | account_alias=C-0B344485; csm=Elena Sinclair; arr=64384.00; trusted_renewal_date=2026-11-16; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=78; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 241 vs prior3 231), seat utilization 78% | account_alias=C-0B369871; csm=Dana Mercer; arr=85128.00; trusted_renewal_date=2026-10-29; trusted_source_why=systems agree (annual term); in_90d_window=True; dates_disagree=False; seat_utilization_pct=75.1; usage_3m_ratio=1.04; risk=low; evidence=3-month usage ratio 1.04 (last3 avg 330 vs prior3 316), seat utilization 75%
arr_at_risk_90d359409.00
disagreementsaccount_alias=C-0B7D2C30; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-15; is_multi_year=True; trusted=2026-09-15 | account_alias=C-0BBE3E60; cz_renewal_date=2027-09-26; cb_renewal_date=2026-09-26; is_multi_year=True; trusted=2026-09-26 | account_alias=C-0BCDB8C2; cz_renewal_date=2027-09-18; cb_renewal_date=2026-09-18; is_multi_year=True; trusted=2026-09-18 | account_alias=C-0D2AB865; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-22; is_multi_year=True; trusted=2026-09-22 | account_alias=C-0F5D2323; cz_renewal_date=2026-09-10; cb_renewal_date=2026-09-29; is_multi_year=True; trusted=2026-09-29
snapshot_date2026-09-05
total_arr_renewing_90d1048715.00
trust_rulemulti-year contracts: Chargebee is authoritative (ChurnZero known wrong); otherwise systems agree or Chargebee wins
window2026-09-05 to 2026-12-04
graded checkstotal_renewing · arr_at_risk · disagreements_flagged · trust_rule

Model's actual answer

show the model's full answer
90-DAY RENEWAL RISK BRIEF
As of 2026-09-24. Window: 2026-09-24 through 2026-12-23. All 20 listed accounts fall in (or before) the window.

DATE-SOURCE DECISION RULE
Chargebee is trusted for all 5 multi-year accounts (ChurnZero is known-wrong on multi-year terms). For the 15 non-multi-year accounts the two systems agree exactly, so no conflict; Chargebee date used for consistency.

DISAGREEMENTS FLAGGED (5 of 20 — all multi-year, all resolved to Chargebee)
- C-0B7D2C30: CZ 2026-09-10 vs CB 2026-09-15 -> use 2026-09-15
- C-0BCDB8C2: CZ 2027-09-18 vs CB 2026-09-18 -> use 2026-09-18 (CZ added a full year)
- C-0D2AB865: CZ 2026-09-10 vs CB 2026-09-22 -> use 2026-09-22
- C-0BBE3E60: CZ 2027-09-26 vs CB 2026-09-26 -> use 2026-09-26 (CZ added a full year)
- C-0F5D2323: CZ 2026-09-10 vs CB 2026-09-29 -> use 2026-09-29

PAST-DUE FLAG: Trusted dates for C-0B7D2C30 (09-15), C-0BCDB8C2 (09-18), and C-0D2AB865 (09-22) are already in the past as of 2026-09-24. The data does not say whether these renewed, lapsed, or are in late renewal — treat as overdue and confirm status immediately.

DEFINITIONS
- Seat utilization = seats_used / seats (from ChurnZero file).
- 3-month trend = active_users change, 2026-06 -> 2026-08 (last 3 reported months).
- Risk rubric: HIGH = sustained usage decline and/or seat utilization under ~35%; MEDIUM = flat usage with moderate (50-70%) utilization; LOW = growing usage and/or utilization above 70%.

RENEWALS (date order, trusted dates)

1. C-0B7D2C30 | Dana Mercer | $65,901 | 2026-09-15 (PAST DUE) | util 274/476 = 57.6% | 3-mo 97->84 (-13.4%)
   HIGH — usage has fallen 12 months straight, 155->84 (-45.8% YoY), and the trusted renewal date has already passed with status unknown.

2. C-0BCDB8C2 | Cole Ingram | $54,427 | 2026-09-18 (PAST DUE) | util 232/424 = 54.7% | 3-mo 127->110 (-13.4%)
   HIGH — steady 12-month decline 200->110 (-45.0%) and renewal date already past.

3. C-0D2AB865 | Elena Sinclair | $38,022 | 2026-09-22 (PAST DUE) | util 250/407 = 61.4% | 3-mo 125->109 (-12.8%)
   HIGH — 12-month decline 199->109 (-45.2%) with renewal date just passed.

4. C-0BBE3E60 | Dana Mercer | $30,993 | 2026-09-26 (2 days) | util 74/114 = 64.9% | 3-mo 39->33 (-15.4%)
   HIGH — steepest decline in the book, 63->33 (-47.6% YoY), renewing in 2 days.

5. C-0F5D2323 | Cole Ingram | $90,647 | 2026-09-29 (5 days) | util 111/390 = 28.5% | 3-mo 20->18 (-10.0%)
   HIGH — only 28.5% of seats used and monthly active users stuck at 17-21 against 390 seats; largest ARR in the cohort.

6. C-0EC6999D | Elena Sinclair | $79,419 | 2026-10-03 (9 days) | util 31/112 = 27.7% | 3-mo 17->15 (-11.8%)
   HIGH — lowest utilization in the book (27.7%) with flat ~15 active users all year, indicating a near-dormant deployment.

7. C-0B20DB64 | Dana Mercer | $21,770 | 2026-10-07 | util 214/378 = 56.6% | 3-mo 294->294 (0.0%)
   MEDIUM — usage perfectly flat at ~294 for 12 months; stable but no growth and 43% of seats idle.

8. C-0BBC4E7A | Cole Ingram | $56,374 | 2026-10-10 | util 228/337 = 67.7% | 3-mo 142->139 (-2.1%)
   MEDIUM — flat usage (~142 all year) at adequate utilization; no decline signal, no expansion signal.

9. C-0FD551AB | Elena Sinclair | $48,815 | 2026-10-14 | util 210/376 = 55.9% | 3-mo 123->126 (+2.4%)
   MEDIUM — flat-to-slightly-up usage but 44% of seats unused.

10. C-0F9F8F13 | Dana Mercer | $46,230 | 2026-10-18 | util 199/352 = 56.5% | 3-mo 185->182 (-1.6%)
    MEDIUM — stable usage (~182-185) with moderate seat waste.

11. C-0BC34584 | Cole Ingram | $16,740 | 2026-10-22 | util 327/494 = 66.2% | 3-mo 104->106 (+1.9%)
    MEDIUM — steady slight upward drift (103->106) at two-thirds utilization.

12. C-0B7A7546 | Elena Sinclair | $35,062 | 2026-10-25 | util 182/205 = 88.8% | 3-mo 64->63 (-1.6%)
    LOW — highest utilization in the book (88.8%) with 12-month usage up 58->63 (+8.6%).

13. C-0B369871 | Dana Mercer | $85,128 | 2026-10-29 | util 317/422 = 75.1% | 3-mo 326->333 (+2.1%)
    LOW — consistent 12-month growth 289->333 (+15.2%) at 75% utilization; expansion candidate.

14. C-0B144C78 | Cole Ingram | $30,899 | 2026-11-02 | util 169/224 = 75.4% | 3-mo 101->106 (+5.0%)
    LOW — usage growing 90->106 (+17.8% YoY) with strong seat adoption.

15. C-0FC4DBB8 | Elena Sinclair | $94,732 | 2026-11-05 | util 356/464 = 76.7% | 3-mo 189->193 (+2.1%)
    LOW — largest ARR in cohort with steady growth 168->193 (+14.9% YoY) and 77% utilization.

16. C-0D5BBE3A | Dana Mercer | $39,740 | 2026-11-09 | util 85/102 = 83.3% | 3-mo 88->91 (+3.4%)
    LOW — growing usage (76->91, +19.7% YoY) at 83% utilization; near seat capacity, upsell candidate.

17. C-0FB9D5AF | Cole Ingram | $63,158 | 2026-11-13 | util 144/199 = 72.4% | 3-mo 173->176 (+1.7%)
    LOW — uninterrupted 12-month growth 154->176 (+14.3%) at healthy utilization.

18. C-0B344485 | Elena Sinclair | $64,384 | 2026-11-16 | util 224/287 = 78.0% | 3-mo 238->244 (+2.5%)
    LOW — steady growth 211->244 (+15.6% YoY) at 78% utilization.

19. C-0CB2C1B4 | Dana Mercer | $40,628 | 2026-11-20 | util 386/473 = 81.6% | 3-mo 47->49 (+4.3%)
    LOW — 81.6% utilization with modest usage growth 43->49 (+14.0% YoY).

20. C-22170CA1 | Cole Ingram | $45,646 | 2026-11-24 | util 251/294 = 85.4% | 3-mo 143->146 (+2.1%)
    LOW — 85.4% utilization and steady growth 130->146 (+12.3% YoY); approaching seat capacity.

TOTALS (arithmetic shown)
Total ARR renewing = 65,901+54,427+38,022+30,993+90,647+79,419+21,770+56,374+48,815+46,230+16,740+35,062+85,128+30,899+94,732+39,740+63,158+64,384+40,628+45,646 = $1,048,715

ARR at risk (HIGH) = 65,901+54,427+38,022+30,993+90,647+79,419 = $359,409 (34.3% of total)
MEDIUM ARR = 21,770+56,374+48,815+46,230+16,740 = $189,929 (18.1%)
LOW ARR = 35,062+85,128+30,899+94,732+39,740+63,158+64,384+40,628+45,646 = $499,377 (47.6%)
Check: 359,409+189,929+499,377 = 1,048,715 ✓

Bottom line: $1,048,715 renews in the window; $359,409 (34.3%) is high risk, concentrated in 6 accounts — 4 with ~45%+ usage declines and 2 with sub-30% seat utilization. Of that, $158,350 (C-0B7D2C30, C-0BCDB8C2, C-0D2AB865) is already past its trusted renewal date with status unknown in the data provided — confirming those three outcomes is the first action. C-0BBE3E60 ($30,993) renews in 2 days and C-0F5D2323 ($90,647) in 5 days.
ticket-theme-synthesis0.80
Plain-English question: What are the real themes in this quarter's support tickets, ignoring the existing tags?
Prompt — what the model saw
Synthesize the quarter's support tickets into 4 to 7 themes. Do not trust existing tags; read the text. Per theme: name, count, share, distinct accounts, ARR affected, two ticket ids, one-line recommendation. Separate broad patterns from single-account noise and rank by ARR exposure, not volume.
cost $0.1386 · 97s · in 11,039 / out 6,272 tok
✗ theme_count_range✓ counts_verifiable✓ single_account_flagged✓ arr_ranked✓ ticket_ids

✓ Correct Answer (ground truth)

Plain-English answer: 80 tickets in 4-7 themes — points-not-posting leads (20 tickets, 25%, $31,100 ARR affected) — each with counts, distinct accounts, ARR, and a recommendation.
The answer in one line: total_tickets = 80; quarter_window = 2026-06-01 through 2026-08-31
total_tickets80
noteexisting tags are noisy/unreliable; classify by text
quarter_window2026-06-01 through 2026-08-31
ranking_by_arr_exposurehris_provisioning, redemption_failures, billing_dispute, points_not_posting, slack_sync
sample_ticket_ids80 items — e.g. IC-460004, IC-460016, IC-460001, IC-460041
single_account_noise_themebilling_dispute
themes{"points_not_posting": {"count": 20, "share_pct": 25.0, "distinct_accounts": 9, "arr_affected": "31100.00", "example_ticket_ids": ["IC-460001", "IC-460002"], "single_account": false}, "slack_sync": {"
graded checkstheme_count_range · counts_verifiable · single_account_flagged · arr_ranked · ticket_ids

Model's actual answer

show the model's full answer
80 tickets, 24 distinct accounts, $284,800 total distinct-account ARR. Theme assignment by body text (tags ignored — e.g. the same billing complaint appears under "billing", "urgent", "how-to", "question", "bug", "feedback"). Each account maps to exactly one theme, so theme ARR sums to the $284,800 total with no double-counting.

Ranked by ARR exposure:

1. HRIS provisioning failures — broad pattern (3 accounts)
   Count: 12 (15.0%) | Distinct accounts: 3 | ARR: $114,000 (C-0DDFC9A7 $48,000 + C-0B2213A9 $36,000 + C-0F6C0F34 $30,000)
   Tickets: IC-460060, IC-460054
   New hires not provisioned; in 3 tickets the provisioning log shows no errors, i.e. silent failure. Highest ARR exposure and concentrated in enterprise accounts.
   Recommendation: Add roster-vs-provisioned reconciliation with alerting, since the sync fails silently.

2. Reward redemption / checkout failures — broad pattern (7 accounts)
   Count: 18 (22.5%) | Distinct accounts: 7 | ARR: $68,800 (11,000 + 10,700 + 10,300 + 9,600 + 9,600 + 8,900 + 8,700)
   Tickets: IC-460025, IC-460032
   Checkout hangs, redemptions fail, gift-card emails/codes never arrive — and in 5 tickets points were deducted despite the failed order.
   Recommendation: Make point deduction transactional with auto-refund on failed fulfillment; the deduct-without-deliver cases are a direct trust/refund liability.

3. Billing seat-count / tier pricing errors — SINGLE-ACCOUNT NOISE
   Count: 16 (20.0%) | Distinct accounts: 1 | ARR: $52,000 (C-0E9C27D1 only)
   Tickets: IC-460069, IC-460078
   All 16 tickets are one account: repeat seat-count errors across at least 3 consecutive invoices plus a renewal charged at the wrong tier. This is 20% of ticket volume but an account-specific billing defect, not a systemic pattern.
   Recommendation: Escalate to a billing owner for a full invoice audit and credit for C-0E9C27D1; treat as one chronic account issue, not a trend.

4. Recognition points not posting — broad pattern (9 accounts, widest account spread)
   Count: 20 (25.0%) | Distinct accounts: 9 | ARR: $31,100 (4,500 + 4,500 + 4,200 + 3,500 + 3,400 + 2,900 + 2,900 + 2,700 + 2,500)
   Tickets: IC-460016, IC-460004
   Recognitions show delivered but points never land; ranges from single-user balances to whole-team failures "after the weekend," suggesting a pipeline/job failure rather than user error.
   Recommendation: Fix the points-ledger pipeline and add delivered-vs-posted reconciliation, with scrutiny on weekend batch jobs.

5. Slack integration failures — broad pattern (4 accounts)
   Count: 14 (17.5%) | Distinct accounts: 4 | ARR: $18,900 (C-10A56B0F $5,400 + C-8C2E8F00 $5,200 + C-0B843542 $4,400 + C-0BA71F12 $3,900)
   Tickets: IC-460047, IC-460046
   Recognition sync stops, the sync toggle resets itself, re-auth does not stick, and the slash command errors for entire teams — all consistent with token/auth-state persistence failure.
   Recommendation: Fix OAuth token refresh and auth-state persistence; add a sync-health check customers can see.

Arithmetic check: 12+18+16+20+14 = 80 tickets; shares 15.0+22.5+20.0+25.0+17.5 = 100%; ARR 114,000+68,800+52,000+31,100+18,900 = $284,800 = sum of all 24 distinct accounts.

Volume vs. exposure note: points-not-posting is the largest theme by volume (25%) but only 4th by ARR ($31.1K, small accounts); HRIS is 3rd by volume but 1st by ARR ($114K). Billing is 20% of volume but one account. Prioritizing by ARR exposure: HRIS > redemption > billing (single account) > points > Slack.
marketing 5 tests
case-study-matchPASS 1.00
Plain-English question: Which three existing customers are the best case-study matches for this prospect?
Prompt — what the model saw
Given the prospect, rank the three most similar existing customers with public case studies for use as social proof. Similarity on industry, size band, use case, region. Cite the fields that drove each match. Do not name any customer without a case study and never include billing data or contact names.
cost $0.0288 · 28s · in 2,316 / out 1,132 tok
✓ top3✓ decoys_excluded✓ cites_fields✓ no_billing

✓ Correct Answer (ground truth)

Plain-English answer: Three ranked matches on industry, size, use case, and region (e.g. C-11C31562), with the higher-scoring decoy explicitly excluded for a mismatched field.
The answer in one line: prospect_alias = C-82AF3719
constraintsno_billing_data, no_contact_names, only_customers_with_has_case_study_true_named
decoys_excludedcompany_alias=C-D6217CAA; similarity_score=4; reason=no_case_study | company_alias=C-C153868F; similarity_score=4; reason=no_case_study
prospect_aliasC-82AF3719
prospect_fields_usedindustry, size_band, use_case, region
top3company_alias=C-11C31562; similarity_score=3 | company_alias=C-64171065; similarity_score=3 | company_alias=C-A13C193D; similarity_score=3
graded checkstop3 · decoys_excluded · cites_fields · no_billing

Model's actual answer

show the model's full answer
Scoring: 1 point per exact match on industry, size_band, use_case, region; max 4. Only customers with has_case_study=true considered.

1. C-64171065 — 3/4 (75%)
   Matches: industry=Technology, size_band=Mid-Market, use_case=employee_recognition
   Differs: region=NA-East vs prospect NA-West

2. C-A13C193D — 3/4 (75%)
   Matches: industry=Technology, size_band=Mid-Market, region=NA-West
   Differs: use_case=retention vs prospect employee_recognition

3. C-11C31562 — 3/4 (75%)
   Matches: size_band=Mid-Market, use_case=employee_recognition, region=NA-West
   Differs: industry=Manufacturing vs prospect Technology

No customer with a public case study matches all 4 fields. Tie-break order: industry+use_case, then industry+region, then use_case+region.
channel-attribution-mathPASS 1.00
Plain-English question: Compute cost per SQM, cost per SQO, and pipeline per dollar for every channel over six months.
Prompt — what the model saw
Compute channel performance for the trailing 6 months. Paid: spend, SQMs, SQOs, cost per SQM, cost per SQO, SQM-to-SQO rate, pipeline amount, pipeline per dollar. Organic: volume, SQO rate, pipeline. Report a channel with spend and zero SQMs as undefined, not zero. Flag rows where SQO date precedes SQM date. Finish with a reallocation recommendation and your confidence given sample sizes.
cost $0.1317 · 118s · in 4,548 / out 7,543 tok
✓ undefined_channel✓ undefined_not_zero✓ flagged_rows✓ best_channel✓ reallocation

✓ Correct Answer (ground truth)

Plain-English answer: LinkedIn Ads at $960/SQM and $3,000/SQO, organic search at $90,000 pipeline on zero spend, paid social flagged at $18,000 with zero SQMs, and a channel winner named.
The answer in one line: best_pipeline_per_dollar_channel = paid_search
best_pipeline_per_dollar_channelpaid_search
confidence_notewebinars and referral have small SQM samples (<20); treat rates as directional.
flagged_rows_sqo_before_sqmCT-000041, CT-000044
months2026-03, 2026-04, 2026-05, 2026-06, 2026-07, 2026-08
per_channel{"linkedin_ads": {"spend_usd": 24000, "sqms": 25, "sqos": 8, "pipeline_amount": 96000, "cost_per_sqm": 960.0, "cost_per_sqo": 3000.0, "sqm_to_sqo_rate": 0.32, "pipeline_per_dollar": 4.0}, "organic_sea
reallocation_recommendationfrom=paid_social; to=paid_search; rationale=paid_social has spend but zero SQMs (cost per SQM undefined, not zero); paid_search has the highest pipeline per dollar.
undefined_channelspaid_social
graded checksundefined_channel · undefined_not_zero · flagged_rows · best_channel · reallocation

Model's actual answer

show the model's full answer
CHANNEL PERFORMANCE — TRAILING 6 MONTHS (2026-03 → 2026-08)
All contacts and spend fall inside the window. 122 contacts total (CT-000001–CT-000122).

PAID CHANNELS
Spend = monthly x 6.

channel       spend      SQMs  SQOs  cost/SQM   cost/SQO   SQM→SQO   pipeline    pipeline/$
paid_search   $36,000    40    18    $900       $2,000     45.0%     $720,000    $20.00
linkedin_ads  $24,000    25    8     $960       $3,000     32.0%     $96,000     $4.00
webinars      $9,000     12    5     $750       $1,800     41.7%     $60,000     $6.67
paid_social   $18,000    0     0     undefined  undefined  undefined $0          $0.00

Arithmetic:
- paid_search: 36,000/40 = 900; 36,000/18 = 2,000; 18/40 = 45.0%; 18 x $40,000 = 720,000; 720,000/36,000 = 20.00
- linkedin_ads: 24,000/25 = 960; 24,000/8 = 3,000; 8/25 = 32.0%; 8 x $12,000 = 96,000; 96,000/24,000 = 4.00
- webinars: 9,000/12 = 750; 9,000/5 = 1,800; 5/12 = 41.7%; 5 x $12,000 = 60,000; 60,000/9,000 = 6.67
- paid_social: $18,000 spend, zero contacts in the file → zero SQMs. Per your rule: cost/SQM, cost/SQO, and SQM→SQO are undefined (division by zero), not zero. Pipeline/$ = 0/18,000 = $0.00.
- Paid blended: spend $87,000; 77 SQMs; 31 SQOs; pipeline $876,000; pipeline/$ = 876,000/87,000 = $10.07.

ORGANIC / NON-PAID (no spend rows)
channel        volume (SQMs)  SQOs  SQO rate  pipeline
organic_search 30             10    33.3%     $90,000
referral       15             6     40.0%     $48,000

- organic_search: 10/30 = 33.3%; 10 x $9,000 = 90,000
- referral: 6/15 = 40.0%; 6 x $8,000 = 48,000

FLAGGED ROWS (SQO date precedes SQM date)
- CT-000044 (linkedin_ads): SQM 2026-07-23, SQO 2026-07-18 — SQO 5 days before SQM
- CT-000041 (linkedin_ads): SQM 2026-06-14, SQO 2026-06-09 — SQO 5 days before SQM
Both are linkedin_ads and both carry $12,000 pipeline. If excluded as invalid: linkedin_ads drops to 6 SQOs, 24.0% rate, $4,000 cost/SQO, $72,000 pipeline, $3.00 pipeline/$. (CT-000007 has same-day SQM/SQO, 2026-03-28 — not a violation.)

REALLOCATION RECOMMENDATION
1. Cut paid_social entirely: $18,000 spent, zero SQMs, zero pipeline. Move the $3,000/mo to paid_search.
2. Shift linkedin_ads budget toward paid_search. paid_search returns $20.00 pipeline per dollar vs linkedin_ads at $4.00 ($3.00 if flagged rows are excluded) — a 5x efficiency gap. paid_search also has the best cost/SQO at scale ($2,000 on 18 SQOs).
3. Keep webinars: $6.67 pipeline/$ beats linkedin_ads on the smallest budget; consider it the secondary paid channel if linkedin_ads is trimmed.
4. organic_search and referral together produced $138,000 pipeline on zero recorded spend — worth protecting, but no spend lever exists in this data.

CONFIDENCE
Low to moderate. Direction is clear but samples are small:
- paid_search is the only channel with a meaningful n (40 SQMs / 18 SQOs); its $20.00 pipeline/$ lead is large enough to act on.
- paid_search 45.0% (n=40) vs linkedin_ads 32.0% (n=25) conversion: pooled rate 26/65 = 40.0%, SE = sqrt(0.40 x 0.60 x (1/40 + 1/25)) = sqrt(0.0156) = 0.125, z = 0.13/0.125 = 1.04 — not statistically significant. The rate gap alone doesn't prove paid_search converts better.
- webinars (12/5) and referral (15/6) are too small for reliable rates.
- 2 of 8 linkedin_ads SQOs (25% of its pipeline) have impossible date ordering — its true performance may be worse than reported.
- paid_social's zero-SQM result on $18,000 spend is the most decision-ready finding: there is no ambiguity in zero.

Caveat: pipeline amounts here are per-contact first-touch values ($40k/$12k/$9k/$8k by channel), so pipeline/$ partly reflects channel-specific deal sizes, not just channel efficiency.
competitor-battlecard-updatePASS 1.00
Plain-English question: Update the Rivally battlecard with sourced pricing and where each side wins.
Prompt — what the model saw
Update the battlecard for the competitor. Sections: one-line positioning, pricing with source and date (newer source wins, note the conflict), where they win, where we win, objections and responses, recent changes, our 12-month win/loss record against them. Cite a snippet id for every factual claim. Rep opinion on a call is not a fact about the competitor. Anything from the old card you cannot re-source gets marked unverified.
cost $0.0583 · 50s · in 2,955 / out 2,971 tok
✓ competitor_named✓ pricing✓ win_loss✓ unverified_marked✓ citations

✓ Correct Answer (ground truth)

Plain-English answer: Rivally at $7 per user per month (the August source wins over the stale $5, conflict noted), a 13-7 win/loss record, and two unverified old claims flagged.
The answer in one line: competitor = Rivally
citation_ids_available25 items — e.g. S01, S02, S03, S04
competitorRivally
pricingvalue_per_user_per_month=7; source_snippet=S17; source_date=2026-08-12
rep_opinion_snippet_ids_not_factsS09, S21
unverified_old_claimsRivally lacks a Slack integration., Rivally was acquired by WorkHuman in 2025.
win_loss_12mowins=13; losses=7
graded checkscompetitor_named · pricing · win_loss · unverified_marked · citations

Model's actual answer

show the model's full answer
# Battlecard: Rivally (updated 2026-09-24)

## One-line positioning
Points-based employee recognition platform with an engaging recognition feed, expanding into engagement surveys (Pulse add-on) and pushing into EU enterprise. [S02, S16, S06, S23, S11, S15]
Note: the old card's "mid-market" framing is not directly supported by any snippet; only a mid-market reviewer exists in the data [S04]. Marked UNVERIFIED.

## Pricing (newer source wins)
- Current list: Recognition Starter at $7 per user/month, annual billing required, per pricing page updated 2026-08-12. [S17]
- Prior price: $5 per user/month, annual billing, per pricing pages dated 2026-01-20 and 2026-04-01. [S03, S08] — superseded by S17.
- CONFLICT / field-reported quotes (not list-price facts):
  - 2026-06-02: prospect reported a $6.50/user/mo quote for 500 seats, annual term — above the then-current $5 list. [S13]
  - 2026-08-14: prospect reported $7/user/mo list with a 15% discount for a 3-year term — consistent with the new $7 list. [S18]
- Pulse survey add-on is priced separately, not bundled (GA 2026-09-01). [S23]

## Where they win
- EU / distributed teams: multi-language support praised by EU enterprise reviewers; EU data residency generally available since 2026-07-01; Dublin office opened. [S12, S15, S11]
- Fast deployment: mid-market reviewer reported setup under a week, Slack integration worked out of the box. [S04]
- End-user engagement: recognition feed repeatedly praised as engaging. [S02, S16]
- Support responsiveness: response time under 4 hours praised. [S22]

## Where we win
- Analytics depth: an 800-seat prospect picked Bonusly over Rivally citing analytics depth [S25]; Rivally's analytics are described as limited [S02], dashboards basic vs. enterprise tools [S07], and exports CSV-only [S20].
- Admin tooling: Rivally's admin tooling lags peers [S16]; admin console lacks bulk recognition editing as of 2026-09-02. [S24]
- Enterprise provisioning: Rivally lacks SCIM; manual user management called painful by an enterprise reviewer. [S10]
- EMEA rewards catalog: thinner than their US catalog per TrustRadius review. [S14]

## Objections and responses
1. "Rivally is cheaper."
   Response: Their list price rose from $5 to $7/user/mo on 2026-08-12, and Pulse is a separately priced add-on, not bundled. [S17, S23]
2. "Rivally covers our EU footprint (data residency, languages)."
   Response: Concede the residency point — it is GA [S15] — then reframe on total capability: their EMEA rewards catalog is thinner than their US one [S14], and analytics/admin gaps remain [S07, S24].
3. "Rivally is faster to deploy."
   Response: Setup under a week is real [S04]; counter with post-launch reality: manual user management without SCIM [S10], CSV-only exports [S20], no bulk editing [S24].
4. "Their feed drives engagement."
   Response: True per reviewers [S02, S16]; pair it with the measurement gap — limited analytics and basic dashboards make ROI hard to prove [S02, S07]; an 800-seat prospect chose us on exactly this. [S25]

## Recent changes (last ~12 months)
- 2025-11-04: $40M Series C led by Northgate Ventures. [S01]
- 2026-03-05: Launched "Rivally Pulse" engagement survey add-on. [S06]
- 2026-05-09: Hired ex-Workday VP EMEA to lead European expansion. [S11]
- 2026-07-01: Dublin office opened; EU data residency generally available. [S15]
- 2026-08-12: List price increased $5 to $7/user/mo. [S17]
- 2026-08-20: Microsoft Teams app v2 in public preview. [S19]
- 2026-09-01: Pulse exits beta; priced as add-on, not bundled. [S23]

## Our 12-month win/loss record vs. Rivally
Window: 2025-09 through 2026-08 (the full deal file, exactly 12 months).
- Total deals: 20. Wins: 13. Losses: 7. Win rate: 13/20 = 65%.
- Losses: Deal-7767F5 (2025-09), Deal-D263E0 (2025-11), Deal-935746 (2025-12), Deal-9066A6 (2026-03), Deal-5645A5 (2026-04), Deal-72A02F (2026-04), Deal-C6FFAA (2026-05).
- Losses cluster in 2026-03 through 2026-05 (4 of 7 losses). Last 3 months on record (2026-06 to 2026-08): 3-0 (Deal-67BE14, Deal-1B6969, Deal-F03E7B).

## Old-card items: verification status
- "Points-based recognition" — VERIFIED [S02, S16]. "Mid-market" qualifier — UNVERIFIED.
- "Starts at $5/user/mo (as of 2026-01)" — OUTDATED; superseded by $7 as of 2026-08-12. [S03, S08, S17]
- "Rivally lacks a Slack integration" — CONTRADICTED by S04 (Slack worked out of the box). Removed.
- "Acquired by WorkHuman in 2025" — UNVERIFIED; no source in the snippet file. Removed pending re-source.
- "Strong in EU enterprise with multi-language support" — VERIFIED [S12, S11, S15].

## Excluded as non-facts (rep opinion, per rules)
- S09 (2026-04-10): AE opinion that Rivally's UI is clunky — unverified against product.
- S21 (2026-08-28): AE opinion that Rivally is discounting aggressively — not confirmed. (S18's prospect-reported 15% 3-year discount is a data point, but one quote does not establish a pattern.)
nurture-sequence-diagnosisPASS 1.00
Plain-English question: What is broken in each email sequence — rates, overlaps, tracking errors?
Prompt — what the model saw
Review the sequences. Per sequence: sent, open rate, reply rate, meeting rate, weakest step. Identify tracking errors (opened above sent), audience overlap between sequences, and the failure mode for anything under 2 percent reply. One change per weak sequence and which to fix first. Under 350 words.
cost $0.3156 · 193s · in 25,250 / out 13,672 tok
✓ tracking_error✓ overlap_pair✓ weakest_step✓ fix_priority

✓ Correct Answer (ground truth)

Plain-English answer: Per-sequence open, reply, and meeting rates with the weakest step (all three weakest at step 3), an opened-above-sent tracking error, and audience overlap flagged.
The answer in one line: fix_priority_first = Cold Outbound - HR Leaders
all_pairwise_overlaps['Cold Outbound - HR Leaders', 'Cold Outbound - People Ops'], ['Expansion Nurture', 'New Logo Nurture']
audience_overlap_pairCold Outbound - HR Leaders, Cold Outbound - People Ops
failure_modeCold Outbound - HR Leaders=opens healthy (~37%) but replies near zero from step 1 -> body/CTA not landing, not a deliverability or subject-line problem
fix_priority_firstCold Outbound - HR Leaders
per_sequence{"New Logo Nurture": {"sent": 1386, "opened": 490, "replied": 90, "meetings": 27, "open_rate": 0.3535, "reply_rate": 0.0649, "meeting_rate": 0.0195, "weakest_step": 3}, "Expansion Nurture": {"sent": 8
tracking_errorsequence=Expansion Nurture; step=2; sent=300; opened=340; issue=opened exceeds sent
under_2pct_replyCold Outbound - HR Leaders
graded checkstracking_error · overlap_pair · weakest_step · fix_priority

Model's actual answer

show the model's full answer
PER SEQUENCE (totals across 3 steps; rates = total / total sent)

New Logo Nurture: sent 1,386 | opened 490 (35.4%) | replied 90 (6.5%) | meetings 27 (1.9%). Weakest step: 3 (reply 18/428 = 4.2%).
Expansion Nurture: sent 875 | opened 565 (64.6%) | replied 59 (6.7%) | meetings 12 (1.4%). Weakest step: 3 (12/275 = 4.4%).
Cold Outbound - HR Leaders: sent 1,785 | opened 545 (30.5%) | replied 8 (0.45%) | meetings 0 (0.0%). Weakest step: 3 (1/590 = 0.17%).
Cold Outbound - People Ops: sent 1,163 | opened 340 (29.2%) | replied 29 (2.5%) | meetings 6 (0.5%). Weakest step: 3 (6/377 = 1.6%).

TRACKING ERROR
Expansion Nurture step 2: opened 340 > sent 300 (113.3%). Open pixel fired more times than sends recorded — bot/re-open counting or data join bug; open rate for this sequence is inflated and unreliable.

AUDIENCE OVERLAP (audiences.csv)
- 21 contacts in both Cold Outbound - HR Leaders and Cold Outbound - People Ops (CT-000849, CT-000884, CT-000890, CT-000908, CT-001033, CT-001097, CT-001101, CT-001103, CT-001105, CT-001130, CT-001153, CT-001159, CT-001217, CT-001227, CT-001236, CT-001255, CT-001258, CT-001277, CT-001285, CT-001311, CT-001345). Same person getting both persona tracks.
- 2 contacts in both New Logo Nurture and Expansion Nurture: CT-000301, CT-000624.

FAILURE MODE (<2% reply)
Cold Outbound - HR Leaders (0.45% reply, 0 meetings): opens are 30.5% but replies collapse across steps (0.83% → 0.34% → 0.17%) — deliverability is fine, message-persona fit is broken (wrong message or wrong titles, not inbox placement). Also 21 of its contacts are simultaneously hit by the People Ops track, diluting attribution.
Note: People Ops step 3 is 1.59% (<2%), sequence total 2.49%.

ONE CHANGE EACH
- HR Leaders: pause and rebuild targeting/message (verify persona, rewrite step 1) — opens prove the email lands; content fails.
- People Ops: dedupe the 21 overlapping contacts into one persona track.
- New Logo Nurture: rework step 3 (biggest drop-off, 8.4% → 4.2%).
- Expansion Nurture: fix step-2 open tracking (340 > 300) before trusting any metric.

FIX FIRST
Cold Outbound - HR Leaders: largest send volume (1,785), 0.45% reply, zero meetings — pure spend with no return, plus it contaminates People Ops results via overlap.
weekly-marketing-goals-rollupPASS 1.00
Plain-English question: Are we on pace on the quarter's marketing goals?
Prompt — what the model saw
Produce the weekly marketing goals update for the active quarter: SQMs, SQOs, DS2s, closed-lost MIA rate, same-quarter close count, active pipeline coverage against target. For each: QTD actual, target, delta, pace (on/behind/ahead given days elapsed). One paragraph on what moved this week.
cost $0.0375 · 32s · in 1,919 / out 1,790 tok
✓ metrics_present✓ deltas✓ coverage_ratio✓ pace✓ paragraph

✓ Correct Answer (ground truth)

Plain-English answer: SQMs ahead of pace, SQOs on pace, DS2s behind (35 short of expected at day 66 of 92), with deltas and same-quarter closes.
The answer in one line: days_in_quarter = 92; days_elapsed = 66; coverage_ratio_active_pipeline = 0.75; elapsed_fraction = 0.7174 (+1 more below)
days_in_quarter92
coverage_ratio_active_pipeline0.75
days_elapsed66
elapsed_fraction0.7174
metrics{"SQMs": {"qtd_actual": 230, "target": 300, "delta": -70, "expected_at_day_elapsed": 215.22, "pace": "ahead"}, "SQOs": {"qtd_actual": 84, "target": 120, "delta": -36, "expected_at_day_elapsed": 86.09,
pace_ruleexpected = target * days_elapsed/days_in_quarter; within +/-5% of expected = on, above = ahead, below = behind; lower_better metrics invert.
this_week_paragraph_required1
graded checksmetrics_present · deltas · coverage_ratio · pace · paragraph

Model's actual answer

show the model's full answer
**Weekly Marketing Goals Update — Q3-2026 (week of day 66/92)**

Pace benchmark: 66/92 days elapsed = **71.7% of quarter**. A metric is "on pace" if QTD attainment ≥ 71.7% of target.

| Metric | QTD Actual | Target | Attainment | Delta | Pace |
|---|---|---|---|---|---|
| SQMs | 230 | 300 | 76.7% | −70 | **Ahead** |
| SQOs | 84 | 120 | 70.0% | −36 | **Behind** (marginal) |
| DS2s | 40 | 75 | 53.3% | −35 | **Behind** |
| Closed-lost MIA rate | 20.0% | ≤10.0% | — | +10 pts | **Off track** |
| Same-quarter closes | 10 | 20 | 50.0% | −10 | **Behind** |
| Active pipeline | $3.0M | $4.0M | 75.0% | −$1.0M | **Behind on coverage** |

Arithmetic:

- **SQMs:** 230/300 = 76.7% vs 71.7% pace → ahead. Linear expectation = 300 × 0.717 = 215.2; actual 230 = +14.8 ahead of pace. Needs 70 more in 26 days (~2.7/day vs 3.5/day QTD run rate — required rate is *lower* than current run rate).
- **SQOs:** 84/120 = 70.0% vs 71.7% → behind by a hair. Expectation = 120 × 0.717 = 86.1; actual 84 = −2.1 off pace. Needs 36 in 26 days (~1.4/day vs 1.27/day QTD run rate — slight acceleration required).
- **DS2s:** 40/75 = 53.3% → meaningfully behind. Expectation = 75 × 0.717 = 53.8; shortfall = −13.8. Needs 35 in 26 days (~1.35/day vs 0.61/day QTD run rate — required rate is 2.2× current).
- **Closed-lost MIA rate:** 5 MIA / 25 closed-lost = 20.0% vs ≤10% target (lower better) → 2× the ceiling. Note: with only 25 closed-lost, each additional MIA loss moves the rate 4 pts; hitting ≤10% at current volume would require ≤2.5 MIA losses — already exceeded.
- **Same-quarter closes:** 10/20 = 50.0% → behind. Expectation = 20 × 0.717 = 14.3; shortfall = −4.3. Needs 10 in 26 days (~0.38/day vs 0.15/day QTD run rate — 2.5× acceleration required).
- **Active pipeline:** $3.0M / $4.0M target = 75% coverage, −$1.0M. Vs linear pace expectation ($4.0M × 0.717 = $2.87M), actual is +$0.13M ahead — but coverage is a snapshot-vs-target metric, so it reads as 25% under target.

**What moved this week:** No prior-week snapshot was provided, so week-over-week movement cannot be quantified from the data given — flagging that explicitly rather than inferring it. What the current snapshot shows: SQMs are the one metric ahead of pace and need only to hold current run rate. SQOs are effectively on pace (−2 vs linear expectation). The structural risks are DS2s and same-quarter closes, both at ~50–53% attainment with 72% of the quarter gone, each requiring a >2× acceleration in weekly run rate to land. The MIA rate is already 2× ceiling and statistically fragile at n=25. Pipeline coverage stands at 75% of the $4.0M target.
communication 5 tests
ceo-slack-compressionPASS 1.00
Plain-English question: Compress the forecast analysis into one Slack message the CEO can read in ten seconds.
Prompt — what the model saw
Compress the forecast analysis below into a single Slack message for the CEO. Rules: at most three sentences; under 80 words; lead with the decision or the so-what; keep every distinct number that matters and the one caveat that changes the conclusion; cut any sentence that does not carry a claim, a number, or a directive. No em dashes, no throat-clearing phrases.
cost $0.0187 · 22s · in 2,201 / out 481 tok
✓ max_3_sentences✓ under_80_words✓ weighted_present✓ caveat_present✓ no_em_dash

✓ Correct Answer (ground truth)

Plain-English answer: Three sentences, under 80 words, decision first, every distinct number kept plus the one caveat — all checked by deterministic text rules.
The answer in one line: weighted_number = 115976; must_keep = data-quality caveat — deleting the bad news fails
weighted_number115,976
limitmax 3 sentences, max 80 words, no em-dashes
must_keepdata-quality caveat — deleting the bad news fails
graded checksmax_3_sentences · under_80_words · weighted_present · caveat_present · no_em_dash

Model's actual answer

show the model's full answer
Q3 weighted forecast is $115,977: $44,729 COMMIT plus 35% of $203,565 BEST_CASE across 54 in-quarter deals. The caveat: 32 deals worth $227,575 slipped to October close dates, including COMMIT Deal-D348E1 ($13,770), which suggests pushes rather than re-qualification. Treat the number as unverified: owner is blank on 85 of 86 deals and none of the 7 COMMIT deals has a documented reason to buy.
stalled-deal-followup-emailPASS 1.00
Plain-English question: Write a follow-up email for this stalled deal — under 120 words, one prior touch, one ask.
Prompt — what the model saw
Write a follow-up email for this stalled deal. Under 120 words. Reference one specific prior touch from the engagement history, make one clear ask, and match the tone of the prior thread. No invented details.
cost $0.0226 · 20s · in 2,025 / out 771 tok
✓ word_cap✓ references_touch✓ single_ask

✓ Correct Answer (ground truth)

Plain-English answer: An email that references the July 28 milestone-awards demo, makes one clear ask, and invents nothing — Deal-0D2F7A, 31 days since last engagement.
The answer in one line: amount = 5100; word_cap = 120; days_since_last_engagement = 31; deal_alias = Deal-0D2F7A (+4 more below)
amount5,100
days_since_last_engagement31
word_cap120
allowed_facts14 items — e.g. Deal-0D2F7A, $5,100, DS3, Alex Franklin
allowed_newENG-00001, ENG-00002, ENG-00003, ENG-00004, ENG-00005
deal_aliasDeal-0D2F7A
must_referenceengagement_id=ENG-00002; date=2026-07-28; type=meeting; why=most substantive prior touch; concrete prospect reactions to milestone awards and points catalog
ownerAlex Franklin
single_cta1
stageDS3
tonewarm, concise, professional (matches prior thread)
graded checksword_cap · references_touch · single_ask

Model's actual answer

show the model's full answer
Note: the data includes no prospect contact name (only "VP People" as a title), so the greeting uses a placeholder. Deal has been stalled 50 days (last engagement 2026-08-05 → today 2026-09-24).

Subject: Re: Pricing recap — 150 seats

Hi [first name],

Following up on my August 5 recap with pricing for 150 seats after your July 28 demo. The People team's reaction to the automated milestone awards and points catalog stood out — and since birthdays and anniversaries are tracked manually today, automating those alone would take real work off your admins' plates.

Would a 20-minute call next week make sense to walk through the pricing and answer any open questions? Happy to work around the VP People's schedule.

Best,
Alex Franklin

Word count: 83 (under 120). Prior touch referenced: ENG-00001 (Aug 5 pricing recap) and ENG-00002 (July 28 demo). Ask: one 20-minute call. Tone matches the prior thread: outbound, concise, recap-style.
funnel-digest-paragraphPASS 1.00
Plain-English question: Write the weekly GTM digest — one paragraph per team, every number traceable.
Prompt — what the model saw
Write the weekly GTM digest: one paragraph per team (Marketing, Sales, CS, PLG) in Bonusly brand voice. Every number must trace to the extracts. No customer contact names, no employee performance commentary, no compensation. Four paragraphs, under 400 words total.
cost $0.0296 · 14s · in 2,133 / out 484 tok
✓ four_paragraphs✓ word_cap✓ teams

✓ Correct Answer (ground truth)

Plain-English answer: Four paragraphs (Marketing, Sales, CS, PLG) using only the allowed numbers, no contact names, no performance commentary, under 400 words.
The answer in one line: word_cap = 400
word_cap400
allowed_numbers16 items — e.g. 2, 3, 9, 08
bannedcustomer contact names, employee performance commentary, compensation
required_paragraphsMarketing, Sales, CS, PLG
graded checksfour_paragraphs · word_cap · teams

Model's actual answer

show the model's full answer
Skill load failed; proceeding with Bonusly brand voice as direct, warm, results-focused, and plainspoken.

Marketing: SQMs came in at 46 against a target of 52 — 88% of goal, six short. The webinar channel drove 18 of those, nearly 40% of weekly volume, while paid search held steady at $150 per SQM. The gap says we need either broader channel mix or stronger conversion on existing spend; webinars are carrying weight that should be shared.

Sales: Fourteen SQOs converted and nine DS2 meetings set, with $310K in new pipeline created. Same-quarter closes landed at three. The conversion-to-meeting ratio (9:14) suggests healthy follow-through, but pipeline depth relative to close count means we need to keep feeding the top or risk a thin quarter-end.

CS: Two renewals saved this week, and team NPS moved to 61. Three red-flag accounts remain open heading into next week. The save rate is solid, but the red-flag count bears watching — retention is won in the margins, and those three represent real risk if they slip.

PLG: 412 new signups with activation at 31%, and 38 companies hit the aha moment of 10 recognition gives. Activation is the metric that matters here; 31% means roughly seven in ten signups aren't reaching value. The 38 companies that did convert are the proof point — the product works when people engage.
partner-digest-sectionPASS 1.00
Plain-English question: Write the partner digest so every partner in the roster appears exactly once.
Prompt — what the model saw
Write the partner digest for the period. Every partner in the roster appears once: active plays get status with numbers, quiet partners get one line. Partner-sourced pipeline count and amount must match the deal data. Nothing omitted, nothing added.
cost $0.0187 · 12s · in 2,071 / out 503 tok
✓ all_partners✓ pipeline_amount✓ quiet_one_line

✓ Correct Answer (ground truth)

Plain-English answer: Four active partners with status and numbers, four quiet partners with one line each, and partner-sourced pipeline that reconciles to the deal data.
The answer in one line: rule = every roster partner named exactly once; quiet partners get one line; no partners outside roster.
active_no_dealsWorkWell Group
partner_sourced_pipelinecount=5; amount=550000
per_partner_pipeline{"Apex Rewards Co": {"count": 2, "amount": 275000}, "HRCloud Partners": {"count": 1, "amount": 140000}, "CultureBridge": {"count": 2, "amount": 135000}}
quiet_partnersRecogniq, KudosWave, PeopleFirst Advisors, TotalPerk
rosterpartner_name=Apex Rewards Co; status=active | partner_name=HRCloud Partners; status=active | partner_name=CultureBridge; status=active | partner_name=WorkWell Group; status=active | partner_name=Recogniq; status=quiet | partner_name=KudosWave; status=quiet … (+2 more)
ruleevery roster partner named exactly once; quiet partners get one line; no partners outside roster.
graded checksall_partners · pipeline_amount · quiet_one_line

Model's actual answer

show the model's full answer
Partner Digest — period 2026-08-24 to 2026-09-04

Active partners

Apex Rewards Co
- 2 sourced opps this period (per Slack 08-25, both UTM Source = Partner).
- Deal-DDAAF2: $180,000; Deal-2CF33E: $95,000. Partner-sourced total: $275,000.
- Co-webinar locked for 09-15. Both opps in DS1 per Slack.

HRCloud Partners
- 1 sourced opp this period: Deal-F1CDA5, $140,000.
- Integration referral cleared security review; opp moved to DS2 (per Slack 08-27).

CultureBridge
- 2 sourced opps from lunch-and-learn (per Slack 08-29), both early stage.
- Deal-096E1D: $60,000; Deal-067213: $75,000. Partner-sourced total: $135,000.

WorkWell Group
- 0 sourced deals this period.
- Wants to restart the joint playbook in Q4; planning call booked for 09-09 (per Slack 09-02).

Quiet partners
- Recogniq: no activity or sourced deals this period.
- KudosWave: no activity or sourced deals this period.
- PeopleFirst Advisors: no activity or sourced deals this period.
- TotalPerk: no activity or sourced deals this period.

Partner-sourced pipeline (from deal data)
- Count: 5 deals (2 Apex + 1 HRCloud + 2 CultureBridge).
- Amount: $275,000 + $140,000 + $135,000 = $550,000.

Note: DS1/DS2 stage labels appear in the Slack export but are not defined in the provided data; reported as stated.
rfp-answer-from-policyPASS 1.00
Plain-English question: Answer the 10 RFP items from the policy excerpts only — no guessing.
Prompt — what the model saw
Answer the 10 questionnaire items from the policy excerpts only. Cite the excerpt id for every answer. Where the excerpts do not answer the question, mark it unanswerable and say what document would; do not guess. Tone: vendor response, direct.
cost $0.0598 · 29s · in 11,921 / out 908 tok
✓ answered_count✓ unanswerable_flagged✓ citations✓ no_guess

✓ Correct Answer (ground truth)

Plain-English answer: Seven answered with excerpt citations (Q1-Q7), three flagged unanswerable with the document that would answer them (Q8-Q10).
The answer in one line: answered_count = 7; rule = unanswerable items must be flagged and the document that would answer them named; do not guess.
answered_count7
answer_map{"Q1": ["P1"], "Q2": ["P1"], "Q3": ["P2"], "Q4": ["P3"], "Q5": ["P4"], "Q6": ["P5"], "Q7": ["P6"]}
ruleunanswerable items must be flagged and the document that would answer them named; do not guess.
unanswerableQ8, Q9, Q10
valid_citation_idsP1, P2, P3, P4, P5, P6, P7, P8
graded checksanswered_count · unanswerable_flagged · citations · no_guess

Model's actual answer

show the model's full answer
Q1: Yes. Customer data is encrypted at rest using AES-256. [P1]

Q2: Yes. Data in transit is encrypted with TLS 1.2 or higher. [P1]

Q3: Yes. Bonusly maintains a SOC 2 Type II report, renewed annually. The report is available to prospects under NDA. [P2]

Q4: Customer data is retained for the duration of the contract plus 90 days, after which it is deleted from production systems. [P3]

Q5: A current list of subprocessors is maintained at trust.bonusly.com. Customers receive 30 days notice before a new subprocessor is added. [P4]

Q6: Yes. SAML 2.0 single sign-on is available on the Enterprise plan. [P5]

Q7: Production databases are backed up daily. Backups are encrypted and retained for 35 days. [P6]

Q8: Unanswerable from the provided excerpts. ISO 27001 certification status would be answered by a current ISO 27001 certificate or by the Trust Center/compliance documentation.

Q9: Unanswerable from the provided excerpts. Uptime SLA terms would be answered by the SaaS Agreement SLA exhibit, customer support SLA documentation, or order form terms.

Q10: Unanswerable from the provided excerpts. HIPAA BAA posture would be answered by the Data Processing Agreement, HIPAA compliance documentation, or legal/security documentation.
ops-maintenance 5 tests
skill-file-reconciliationPASS 1.00
Plain-English question: Audit a skill set against its manifest and flag every conflict with a severity and an action.
Prompt — what the model saw
Reconcile the skill set below against its manifest. Find and report (with severity CRITICAL/WARNING/INFO and an action type TRIM_DESC, UPDATE_BODY, DELETE_SKILL, MERGE, or REVIEW): (1) any two skills whose ALWAYS-trigger phrases overlap/duplicate, (2) any circular delegation chain between skills (name it), (3) any delegation target that does not exist (dangling, name it), (4) any version conflict and which skill should survive, (5) how many manifest descriptions exceed 1,024 characters, (6) any hardcoded page ids, dates, or person names in skill bodies, (7) manifest drift in BOTH directions (files with no manifest row, and manifest rows with no file). Make one proposal per finding; rewrite nothing.
cost $0.3813 · 177s · in 66,421 / out 11,807 tok
✓ duplicate_trigger_found✓ circular_chain✓ dangling_target✓ oversized_desc_count✓ hardcoded_values✓ manifest_drift✓ severity_and_action

✓ Correct Answer (ground truth)

Plain-English answer: A severity-ranked list of TRIM_DESC / UPDATE_BODY / DELETE_SKILL / MERGE / REVIEW actions covering every rule violation found.
The answer in one line: files_audited = 14; rank = CRITICAL / WARNING / INFO with one fix action per issue
files_audited14
findingsduplicate trigger, circular delegation chain, dangling file reference, oversized description, hardcoded values, manifest drift
rankCRITICAL / WARNING / INFO with one fix action per issue
graded checksduplicate_trigger_found · circular_chain · dangling_target · oversized_desc_count · hardcoded_values · manifest_drift · severity_and_action

Model's actual answer

show the model's full answer
SKILL SET RECONCILIATION — 14 files vs skill_manifest.csv
=============================================================

(1) ALWAYS-TRIGGER OVERLAP
--------------------------
Finding 1a — CRITICAL — ACTION: MERGE
Pair: comms-drafter ⇄ email-drafter
Exact duplicate trigger phrases in both descriptions (5):
  "write me an email"
  "draft a follow-up"
  "what should I say"
  "bump email"
  "contract nudge"
Near-duplicate: "help me reply" (comms-drafter) vs "help me reply to this" (email-drafter).
Both skills also claim the same scope: "whenever anyone asks you to write, draft, review, or improve" customer-facing email. Neither contains a lane marker referencing the other (each only disambiguates against deal-strategy-coach), so both will double-fire on every email request.
Proposal: MERGE. comms-drafter survives (superset: sales + CS + support/Intercom + partner comms). Fold email-drafter's two unique blocks into it: the Gmail signature-retrieval procedure and the HubSpot→Granola→Gong transcript source priority. Delete email-drafter after merge.

Finding 1b — WARNING — ACTION: TRIM_DESC
Pair: weekly-pipeline-report ⇄ pipeline-intelligence-report
  weekly-pipeline-report triggers: "run the pipeline update", "do the pipeline report", "update the pipeline", "what does pipeline look like"
  pipeline-intelligence-report triggers: "pipeline update", "run the pipeline report", "what's the pipeline look like"
"pipeline update" is a verbatim substring of "run the pipeline update"; the other two pairs are near-duplicates. Note next-to-close already disambiguates itself against pipeline-intelligence-report in its description; these two do not disambiguate against each other.
Proposal: TRIM_DESC on both. weekly-pipeline-report keeps scheduled-cadence phrasing ("weekly pipeline report", "this week's numbers", "mid-month pipeline check"); pipeline-intelligence-report keeps scored/tiered phrasing ("score the pipeline", "pipeline intelligence", "full pipeline"). Remove "pipeline update" from one of them.

(2) CIRCULAR DELEGATION CHAIN
-----------------------------
Finding — WARNING — ACTION: REVIEW
Cycle: deal-strategy-coach → email-drafter → deal-strategy-coach
  deal-strategy-coach body: "When drafting manager-to-prospect emails, use the email-drafter skill"
  email-drafter body: "if the user needs strategic deal coaching… point them to the deal-strategy-coach skill… suggest they use deal-strategy-coach for the deeper analysis"
Mutual delegation with no explicit termination condition; comms-drafter → deal-strategy-coach ("For deep deal strategy, use deal-strategy-coach") makes it a three-node loop entry point.
Checked and ruled out: pipeline-intelligence-report → closed-lost-analysis is one-directional (closed-lost-analysis Mode 4 "called from pipeline-intelligence-report" is an inbound annotation, not a delegation edge). next-to-close → pipeline-intelligence-report has no return edge. sales-forecast → analysis-validator, claim-compressor → analysis-validator, signalforge-feedback → {analysis-validator, claim-compressor} are sequencing, not cycles.
Proposal: REVIEW. Make the handoff one-directional per task type: strategy requests route deal-strategy-coach → email-drafter for the draft only; email-drafter drops its "suggest deal-strategy-coach" line when invoked via that handoff.

(3) DANGLING DELEGATION TARGETS (named; "exists" = present in the provided 14-file set/manifest)
----------------------------------------------------------------------------------------------
Finding — CRITICAL — ACTION: REVIEW
  a. prospect-research-multithreading — hard "invoke" from comms-drafter, deal-strategy-coach, email-drafter. Not in set.
  b. bonusly-brand — mandatory step ("apply the bonusly-brand skill") in comms-drafter (Step 0), email-drafter, sales-forecast; referenced by signalforge-claim-compressor. Not in set.
  c. signalforge-reports — "MANDATORY PRE-BUILD STEPS" reads in pipeline-intelligence-report; required reads in weekly-pipeline-report. Not in set.
Finding — WARNING — ACTION: REVIEW
  d. analysis-validator §12.4 delegates to 8 targets, none in set: bonusly-data-questions, bonusly-product-questions, bonusly-business-reporting-questions, bonusly-rewards-questions, bonusly-ppp-questions, bonusly-feature-flag-questions, bonusly-deal-desk-questions, bonusly-datadog-questions.
  e. skill-orchestrator — referenced by signalforge-feedback (activation checklist) and analysis-validator §11 (Three-Way Sync cascading files). Not in set.
Finding — INFO — ACTION: REVIEW
  f. caveman — signalforge-claim-compressor ("Both can be used together"). Not in set.
  g. xlsx public skill (scripts/recalc.py path) — stale-pipeline-report Phase 7 validation. Not in set.
Scope caveat: bodies reference /mnt/skills/organization/ and /mnt/skills/user/ paths, so some of these may be org-level skills intentionally outside this manifest. On the data provided they are dangling; if they are in-scope elsewhere, downgrade a–e to INFO and annotate.
Proposal: REVIEW. Either register all 13 targets as manifest rows or add an explicit external-dependency annotation per referencing skill so the manifest is self-describing.

(4) VERSION CONFLICT
--------------------
Finding — WARNING — ACTION: UPDATE_BODY
analysis-validator conflicts with itself:
  Header block: "Version: 3.6", "Last Updated: May 9, 2026 (v3.6…)"; changelog latest = 3.6 (May 9, 2026); footer = "analysis-validator v3.6 · May 9, 2026"
  Section 7 Validation Trail template literal: "Validator: analysis-validator v3.2"
Survivor: v3.6 (header, changelog, and footer all agree; v3.2 appears once, inside the output template).
Proposal: UPDATE_BODY. Change the Section 7 trail template to v3.6 (or to a non-versioned placeholder) so the template stops stamping a stale version into every validated report. Note pipeline-intelligence-report's footer already hardcodes "Analysis Validator v3.6", consistent with the survivor.
Adjacent observation (INFO): pipeline-intelligence-report is the only skill using a non-semver version string ("v6 · May 2026"); the rest use semver (3.6, 1.1, 1.0). Not a conflict, but a scheme inconsistency.

(5) DESCRIPTIONS EXCEEDING 1,024 CHARS
--------------------------------------
Finding — INFO — ACTION: TRIM_DESC (preventive)
Arithmetic over description_chars column:
  656, 897, 996, 792, 965, 676, 945, 1004, 1006, 962, 1006, 708, 762, 656
  max = 1006; 1006 ≤ 1024 → count exceeding = 0
Zero descriptions exceed 1,024. Three are within 20 chars of the cap:
  pipeline-intelligence-report 1006 (18 under)
  signalforge-claim-compressor 1006 (18 under)
  partner-digest 1004 (20 under)
Proposal: TRIM_DESC on those three before the next description edit pushes them over; Finding 1b's trim of pipeline-intelligence-report would also relieve the tightest case.

(6) HARDCODED PAGE IDs, DATES, PERSON NAMES IN BODIES
-----------------------------------------------------
Finding — WARNING — ACTION: UPDATE_BODY
Instances by skill (aliases cited exactly as written):
  partner-digest — Cloud ID 73fe98de-a4a3-4869-9f8a-bb1eeed4cf7f; Space ID 1958248479; folder ID 2286616609; page IDs 2286321666, 2265382925, 2236940297, 2237825028, 2239365136, 2238283777; Slack user ID U03QLMBL7AR; person names Amani Phipps, Kelli, Jen Lee, Hani, Bryce, Sara; date "May 16, 2026 issue".
  sales-forecast — Space ID 2232811524; parent page ID 2232582148; cloud ID; person names Alaina (Step 2A), Elena (changelog); quarter hardcodes despite v1.1 "quarter-agnostic" claim: Step 1A "Open Q2 Deals", Tab 6 "Q2 Narrative", title example "July 9, 2026".
  signalforge-feedback — Page ID 2295136266; spaceId 2232811524; parent 2234417154; Build Log page ID 2247295002; example names "Gavin Porter Rep Diagnostic", "Lowe's Conversation Analysis", "Q2 Pipeline Review".
  pipeline-intelligence-report — hardcoded AE roster "verified May 2026": Bryce Harmon 119337721, Dana Mercer 83155923, Cole Ingram 83155924, Alex Franklin 84342457, Gavin Porter 1520255671; HubSpot org ID 1973303; stage IDs; "last modified March 2023".
  analysis-validator — §12.3 GTM roster with 19 named people + owner IDs (Alaina Loori, Shealagh Coughlin, Core 6 AEs incl. Hugo Lindqvist 77260721, 7 CSMs, Ben Castelli, Amani Phipps, John Thomas, Yasmin Wahid); Finance escalation "Manish or Amani"; dates (March 28, 2023; May 4, 2026; "as of May 2026" ranges); stage IDs 150582536–150582539, 1175632767.
  weekly-pipeline-report — person name Ben Lavin (title + "present the file to Ben"); spreadsheet IDs 1CLZeOsElVDF_LF0ZG_t2nfwvhnZ6bpwqM_nX3WEYzcw and 1ENuaEcCuLjdKhMvp8FK3Ys1ek5Aw9ZuOZhsHJJFoB_k; static Q1 2026 actuals ($365,152 / $475,000; $2,490,532 / $3,288,000); Q2 window "April 1 – June 30, 2026".
  stale-pipeline-report — Slack channel ID C0561C1JCPJ; owner ID 55483190 hardcoded as an exclusion — direct self-contradiction of its own Phase 2 rule "Never hardcode rep names or owner IDs"; "Alaina" in description; "all 97 deals" hardcoded count; example dates 5/7, 5/15, 5/19.
  deal-strategy-coach — Confluence page ID 2257879045 (AE Excellence Playbook, "April 2026"); routing names "routed to Perseus" (India), "routed to Farid" (.edu); "200+ Gong calls and 370+ resolved deals".
  model-selection — registry dates by design: last_checked 2026-05-19, model ID claude-haiku-4-5-20251001, "deprecation announced April 14, 2026".
  closed-lost-analysis — named companies with dates as case examples: MinIO (May 4–12), Estee Lauder, Softheon (May 2026), LIFTOFF, Nestlé, Ozinga, Aurora Innovation, GCash, Ethos Cannabis, StickerYou; "30-deal AI-field sample from May 2026: 10 of 10".
  signalforge-claim-compressor — example entities Felix Construction ($15K TCV), Panopto, Schneider Downs; "Forked… from JuliusBrussee/caveman"; date 2026-05-09.
  next-to-close — HubSpot org ID 1973303 in URL pattern; stage IDs.
  Clean: comms-drafter, email-drafter (no page IDs, dates, or person names found).
Observed drift already caused by this (evidence, not hypothesis): analysis-validator §12.3 defines "Core 6 AEs" including Hugo Lindqvist (77260721); pipeline-intelligence-report's hardcoded AE list omits him — 6 vs 5. Count: 6 − 5 = 1 roster member lost to hardcode drift.
Proposal: UPDATE_BODY. Externalize rosters, page/folder IDs, and static actuals into shared reference blocks resolved at runtime (stale-pipeline-report Phase 2 already demonstrates the pattern with dynamic owner resolution); keep named companies only where explicitly labeled as illustrative examples.

(7) MANIFEST DRIFT — BOTH DIRECTIONS
------------------------------------
Finding — INFO — no action
Arithmetic:
  Files provided: 14 (analysis-validator, closed-lost-analysis, comms-drafter, deal-strategy-coach, email-drafter, model-selection, next-to-close, partner-digest, pipeline-intelligence-report, sales-forecast, signalforge-claim-compressor, signalforge-feedback, stale-pipeline-report, weekly-pipeline-report)
  Manifest rows: 14 (same 14 names)
  Files with no manifest row: 14 − 14 = 0
  Manifest rows with no file: 14 − 14 = 0
  declared_name vs frontmatter name mismatches: 0/14
No drift in either direction.

SUMMARY
-------
CRITICAL: 2 (1a merge comms-drafter/email-drafter; 3 dangling hard-dependency targets)
WARNING: 4 (1b trigger trim; 2 delegation cycle; 4 version template; 6 hardcoded IDs/names/dates)
INFO: 2 (5 zero over 1,024, three near cap; 7 no drift)
No rewrites performed — proposals only, as requested.
runbook-from-slack-threadPASS 1.00
Plain-English question: Turn the resolved incident thread into a runbook.
Prompt — what the model saw
Turn the resolved incident thread into a runbook: ordered steps with the exact command or action taken, who ran it, how success was verified, and rollback for each step that changed state. Every step must trace to a message in the thread. Add nothing the thread does not contain; where the thread is unclear, mark the step as needing confirmation.
cost $0.0467 · 38s · in 3,138 / out 1,788 tok
✓ ordered_steps✓ rollback✓ needs_confirmation✓ sourced

✓ Correct Answer (ground truth)

Plain-English answer: Ordered steps from the thread, each with the exact command, who ran it, how success was verified, and rollback where state changed.
The answer in one line: rule = no steps beyond what the thread contains; every step traces to a message id.
ambiguous_steporder=4; msg=M05; mark=needs-confirmation
ruleno steps beyond what the thread contains; every step traces to a message id.
stepsorder=1; msg=M01; actor=Bryce Harmon; action=acknowledge PagerDuty alert, take IC; state_change=False; verified_by=M01 | order=2; msg=M02; actor=Farid Osman; action=check queue depth; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M02 (48,213 pending) | order=3; msg=M04; actor=Farid Osman; action=pause auto-enqueue via feature flag; command=bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'; state_change=True; rollback=bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'; verified_by=M07 (queue depth falling) | order=4; msg=M05; actor=Elena Sinclair; action=cleared the dead set; state_change=True; needs_confirmation=thread gives no command or method for clearing the dead set | order=5; msg=M06; actor=Bryce Harmon; action=scale reward-worker to 6 replicas; command=kubectl scale deployment/reward-worker --replicas=6; state_change=True; rollback=kubectl scale deployment/reward-worker --replicas=3; verified_by=M07/M08 | order=6; msg=M08; actor=Cole Ingram; action=verify queue drained and error rate baseline; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M08 (returns 0) … (+2 more)
graded checksordered_steps · rollback · needs_confirmation · sourced

Model's actual answer

show the model's full answer
RUNBOOK — reward-worker queue backlog (source: #incident-2026-08-29-reward-queue, resolved 2026-08-29)

Note: every step cites its source message. Gaps are marked NEEDS CONFIRMATION.

Step 1 — Acknowledge and take IC [M01]
- Actor: Bryce Harmon
- Action: Acknowledge PagerDuty alert (reward-worker queue depth > 10k); take incident commander role.
- State change: No. Rollback: N/A.
- Verification: None documented.

Step 2 — Measure queue depth [M02]
- Actor: Farid Osman
- Command: bundle exec rake sidekiq:queue_depth
- Result: 48,213 pending jobs (normal: under 500).
- State change: No. Rollback: N/A.
- Verification: Command output is the verification.

Step 3 — Inspect dead set [M03]
- Actor: Farid Osman
- Action: Checked dead set; found 112 jobs, all Redis::TimeoutError from around 13:58.
- Command: Not documented — NEEDS CONFIRMATION.
- State change: No. Rollback: N/A.
- Verification: Reported finding only.

Step 4 — Pause enqueue to stop the bleed [M04]
- Actor: Farid Osman
- Command: bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'
- State change: Yes (feature flag disabled).
- Rollback (documented in M04): bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'
- Verification: No direct flag-state verification recorded — NEEDS CONFIRMATION. Later queue improvement (M07/M08) is not isolated proof of this step's effect.

Step 5 — Clear the dead set [M05]
- Actor: Elena Sinclair
- Action: "While I was in the console I cleared out the dead set."
- Exact command: Not documented — NEEDS CONFIRMATION.
- State change: Yes (destructive — 112 dead jobs removed).
- Rollback: Not documented — NEEDS CONFIRMATION.
- Verification: None documented — NEEDS CONFIRMATION.
- Caution: Do not treat this as an approved repeatable remediation; the thread only records that it happened.

Step 6 — Scale workers up [M06]
- Actor: Bryce Harmon
- Command: kubectl scale deployment/reward-worker --replicas=6 (was 3)
- State change: Yes (3 to 6 replicas).
- Rollback (documented in M06): kubectl scale deployment/reward-worker --replicas=3
- Verification: No direct replica-count verification recorded — NEEDS CONFIRMATION. M07/M08 are queue observations, not isolated proof of this action's effect.

Step 7 — Monitor drain rate [M07]
- Actor: Farid Osman
- Action: Reported queue depth 9,400 and falling ~1,200/min.
- Measurement command: Not documented — NEEDS CONFIRMATION.
- State change: No. Rollback: N/A.

Step 8 — Verify queue drained [M08]
- Actor: Cole Ingram
- Command: bundle exec rake sidekiq:queue_depth
- Result: Returns 0; Datadog error rate back to baseline (reported).
- State change: No. Rollback: N/A.
- Verification: Command output is the verification.

Step 9 — Re-enable enqueue [M09]
- Actor: Bryce Harmon
- Command: bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'
- State change: Yes (feature flag re-enabled).
- Rollback: Not documented — NEEDS CONFIRMATION. (Disabling the flag was the rollback for Step 4; the thread does not designate it as the approved rollback for this step.)
- Verification: 40 new jobs processed cleanly in the next 3 minutes (reported by Bryce Harmon).

Step 10 — Scale workers back down [M10]
- Actor: Bryce Harmon
- Command: kubectl scale deployment/reward-worker --replicas=3
- State change: Yes (6 to 3 replicas).
- Rollback: Not documented — NEEDS CONFIRMATION.
- Verification: Queue stable at 0; incident declared resolved (reported by Bryce Harmon).

Open items needing confirmation before this runbook is reused:
1. Step 3: exact dead-set inspection command.
2. Step 4: how to verify the feature flag state directly.
3. Step 5: exact dead-set clear command, its verification, and any rollback.
4. Step 6: direct replica-count verification (e.g., a kubectl get command — not in thread).
5. Step 7: measurement command for drain rate.
6. Steps 9 and 10: documented rollbacks.
datadog-incident-timelinePASS 1.00
Plain-English question: Sequence the failure from the logs — what broke first, and what cascaded after it?
Prompt — what the model saw
From the log slice, sequence the failure: first error with timestamp, the cascade in order, the service and job involved, and the Datadog query that would confirm the first error. State what the logs do not show.
cost $0.0507 · 42s · in 3,258 / out 2,398 tok
✓ first_error_ts✓ service_named✓ cascade_ordered✓ query✓ gaps_stated

✓ Correct Answer (ground truth)

Plain-English answer: A reward-service Redis timeout at 14:01:12Z cascades to sidekiq at 14:01:40 and api-gateway at 14:03:05, with the Datadog query that confirms the first error.
The answer in one line: job_class_involved = RewardGiveJob
cascade_orderservice=reward-service; first_error_at=2026-09-03T14:01:12Z | service=sidekiq; first_error_at=2026-09-03T14:01:40Z | service=api-gateway; first_error_at=2026-09-03T14:03:05Z | service=web-app; first_error_at=2026-09-03T14:03:30Z
clean_servicespostgres
confirming_queryservice:reward-service status:error "@message:Redis::TimeoutError"
first_errortimestamp=2026-09-03T14:01:12Z; service=reward-service; message=Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s
job_class_involvedRewardGiveJob
logs_do_not_showroot cause of the Redis timeout, redis-server's own logs, infrastructure metrics (CPU/memory/network), any deploy or config change preceding the incident, who applied mitigation or when (no actor fields in logs)
graded checksfirst_error_ts · service_named · cascade_ordered · query · gaps_stated

Model's actual answer

show the model's full answer
FAILURE SEQUENCE (from datadog_logs.csv only)

1) First error
   2026-09-03T14:01:12Z — service: reward-service — ERROR — "Redis::TimeoutError: Connection to Redis at redis-primary:6379 timed out after 5s"
   (All prior entries, 13:55:13Z–13:59:30Z, are INFO: health checks, cache hits, job enqueues. Nothing anomalous before 14:01:12Z.)

2) Cascade, in order
   - 14:01:12Z  reward-service  Redis connection timeout (first error)
   - 14:01:20Z / 14:01:30Z / 14:01:40Z  reward-service  "retry exhausted for RewardGiveJob" x3
   - 14:01:40Z  sidekiq  "RewardGiveJob failed: Redis::TimeoutError; retrying in 60s" (first worker-side failure)
   - 14:02:28Z  sidekiq  RewardGiveJob failed; retrying
   - 14:02:30Z  sidekiq  WARN "Queue reward depth above 10,000" (backlog signal, ~50s after first error)
   - 14:03:05Z  api-gateway  "502 upstream timeout calling reward-service /gives" (user-facing impact begins ~1m53s after first error)
   - 14:03:30Z  web-app  "Give form submission failed: upstream 502 from api-gateway"
   - 14:03:31Z–14:06:52Z  repeating loop: sidekiq retries (14:03:31, 14:04:22, 14:05:26, 14:06:47), api-gateway 502s (14:03:48, 14:04:13, 14:05:16, 14:06:52), web-app form failures (14:04:45, 14:05:42, 14:06:49)
   - 14:22:10Z  reward-service  INFO "Redis connection restored; resuming job processing"
   - 14:24:45Z  sidekiq  INFO "Queue reward depth below 500" (drain/recovery)

3) Service and job
   - Originating service: reward-service (dependency: Redis at redis-primary:6379)
   - Job: RewardGiveJob (log slice). sidekiq_jobs.csv also shows RecognitionDigestJob failing (J-00013–J-00016, 14:02:36Z–14:05:50Z, same Redis::TimeoutError) — those jobs are NOT in the log slice.

4) Datadog query to confirm the first error
   service:reward-service status:error "Redis::TimeoutError"
   Time window: 2026-09-03 13:55:00Z – 14:05:00Z, sort by timestamp ascending, limit 1 — expected top hit: 14:01:12Z "Connection to Redis at redis-primary:6379 timed out after 5s".
   Tighter variant: service:reward-service status:error "Connection to Redis at redis-primary:6379 timed out"

ARITHMETIC
- Error volume in slice: 19 ERROR lines = reward-service 4 (14:01:12, :20, :30, :40) + sidekiq 6 (14:01:40, 14:02:28, 14:03:31, 14:04:22, 14:05:26, 14:06:47) + api-gateway 5 (14:03:05, 14:03:48, 14:04:13, 14:05:16, 14:06:52) + web-app 4 (14:03:30, 14:04:45, 14:05:42, 14:06:49); plus 1 WARN (queue depth).
- sidekiq_jobs.csv: 16 failed jobs = 12 RewardGiveJob + 4 RecognitionDigestJob, 100% Redis::TimeoutError. Earliest failed_at: J-00005 at 14:01:46Z (34s after first log error).
- Outage window: 14:01:12Z → 14:22:10Z restore = 20m58s; → 14:24:45Z queue drained = 23m33s total.
- Time to user impact: 14:01:12Z → 14:03:05Z first 502 = 1m53s.

WHAT THE LOGS DO NOT SHOW
- Root cause of the Redis timeout: no redis-primary entries exist in the slice; nothing shows whether it was network, failover, restart, or resource exhaustion.
- The 14:06:52Z–14:22:10Z gap: no reward-service/sidekiq/api-gateway/web-app entries for ~15 minutes; the slice only shows routine postgres checkpoints. What happened between last retry and "connection restored" is unobserved here.
- RecognitionDigestJob failures (in sidekiq_jobs.csv but zero log lines).
- Any mapping between job_ids (J-00001–J-00016) and log lines — no correlation IDs; per-job outcomes (retried to success vs. dead-set) cannot be confirmed.
- Whether any jobs were permanently lost; logs only say "retrying," never final disposition.
- Actual queue depth numbers (only "above 10,000" and later "below 500"), total affected users/submissions beyond the 4 logged web-app failures, and what action restored Redis.
feature-flag-state-summaryPASS 1.00
Plain-English question: Summarize what each feature flag controls and who has it on.
Prompt — what the model saw
Summarize the state of the listed feature flags: what each controls per the code excerpt, which companies or segments have it on, targeting rules, and any flag with no code reference or no targets. Company counts must match the export. Do not describe a flag that is not in the export.
cost $0.0507 · 38s · in 3,565 / out 1,971 tok
✓ orphans✓ no_targets✓ counts_match✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: Recognition streaks on for 42 beta companies, redeem-flow redesign off with a 12-company target list, plus any flag with no code reference or no targets.
The answer in one line: rule = company counts must equal the export; do not describe flags absent from the export.
code_referencedrecognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, redeem_flow_redesign, analytics_dashboard_v3, ms_teams_app_v2
flags{"recognition_streaks_v2": {"state": "on", "targeting_rules": "segment:beta_companies", "company_count": 42}, "points_budget_guardrails": {"state": "on", "targeting_rules": "all_companies", "company_c
flags_with_no_targetspaused_offboard_cleanup
orphan_flags_no_code_referencelegacy_give_modal, survey_boosters_q3, paused_offboard_cleanup
rulecompany counts must equal the export; do not describe flags absent from the export.
graded checksorphans · no_targets · counts_match · no_invented

Model's actual answer

show the model's full answer
FEATURE FLAG STATE SUMMARY
==========================

Flags in export: 9 (5 on, 4 off). Company counts taken verbatim from flags_export.csv. No company names or IDs were provided, so "who has it on" is reported as segment/rule + count only — no company names can be cited without inventing data.

1. recognition_streaks_v2 — ON
   - Controls (app/models/recognition.rb): when enabled, calls StreakTracker.record(give) on a recognition give — i.e., streak tracking on recognition activity.
   - Targeting: segment:beta_companies — 42 companies.

2. points_budget_guardrails — ON
   - Controls (app/services/budget_service.rb): when enabled, BudgetService.new(company).enforce!(giver, points) runs — enforces point-giving budget limits.
   - Targeting: all_companies — 220 companies. This is the only flag rolled out to all companies.

3. slack_dm_nudges — ON
   - Controls (app/jobs/nudge_job.rb): gate guard — job returns early unless flag is on; when on, SlackDm.send_nudge(user) sends a Slack DM nudge.
   - Targeting: segment:region_na (North America region) — 87 companies.

4. redeem_flow_redesign — OFF
   - Controls (app/controllers/redeem_controller.rb): branch — on renders RedeemV2Component, off renders RedeemV1Component. Currently OFF, so targeted companies still get the V1 redeem flow.
   - Targeting: targeted_list (explicit company list) — 12 companies.
   - Note: flag is off but has a targeted list configured — 12 companies are staged for the V2 rollout when the flag flips.

5. analytics_dashboard_v3 — ON
   - Controls (app/controllers/analytics_controller.rb): when enabled, @dashboard = AnalyticsV3.new(company) — serves the V3 analytics dashboard.
   - Targeting: segment:tier_three — 65 companies.

6. ms_teams_app_v2 — OFF
   - Controls (app/services/teams_installer.rb): when enabled, TeamsAppV2.install(company) — installs the V2 MS Teams app.
   - Targeting: targeted_list — 9 companies.

7. legacy_give_modal — OFF
   - No code reference in flag_code.md — what it controls cannot be determined from the provided excerpt.
   - Targeting: segment:legacy_plan — 14 companies.

8. survey_boosters_q3 — ON
   - No code reference in flag_code.md — what it controls cannot be determined from the provided excerpt.
   - Targeting: segment:legacy_plan — 7 companies.

9. paused_offboard_cleanup — OFF
   - No code reference in flag_code.md — what it controls cannot be determined.
   - Targeting: none — targeting_rules field is empty, company_count = 0. This is the only flag with no targets at all.

Cross-flag observations
- Company-count arithmetic: 42 + 220 + 87 + 12 + 65 + 9 + 14 + 7 + 0 = 456 total flag-company assignments across 9 flags (companies can appear in multiple flags, so this is not a unique-company count — the export gives no unique-company total).
- Segment overlaps visible in the export: legacy_plan is targeted by two flags (legacy_give_modal: 14, survey_boosters_q3: 7). Whether those 7 are a subset of the 14 cannot be determined from the export.
- Flags with no code reference in the excerpt: legacy_give_modal, survey_boosters_q3, paused_offboard_cleanup.
- Flags with no targets: paused_offboard_cleanup only (empty rule, 0 companies).
- points_budget_guardrails is the only all-companies rollout (220); everything else is segment- or list-scoped.
nda-triagePASS 1.00
Plain-English question: Triage each NDA green, yellow, or red — and quote the clause that decides it.
Prompt — what the model saw
Triage each NDA as GREEN (standard approval), YELLOW (counsel review), or RED (full legal review). Cite the clause that drives each classification by quoting it. Catch embedded non-solicits, non-competes, missing carve-outs, and unusual term or governing law. One line of reasoning per NDA.
cost $0.0575 · 21s · in 12,007 / out 735 tok
✓ green✓ yellow✓ red✓ non_solicit_caught✓ clause_quoted✓ correct_labels

✓ Correct Answer (ground truth)

Plain-English answer: NDA-1 green (standard mutual terms), NDA-2 yellow for an embedded 18-month non-solicit, with the driving clause quoted for every classification.
The answer in one line: non_solicit_in = NDA-2
checksclassification_per_nda, clause_quotes_exist_in_document, non_solicit_caught, carveouts_checked
classifications{"NDA-1": {"classification": "GREEN", "driving_reason": "standard mutual NDA: 2-year term, standard carve-outs, Delaware law", "driving_clause_quote": "Carve-outs: information that (a) is or becomes p
non_solicit_inNDA-2
graded checksgreen · yellow · red · non_solicit_caught · clause_quoted · correct_labels

Model's actual answer

show the model's full answer
NDA TRIAGE — three inbound NDAs (party names as given)

NDA-1 — [PARTY A] and Bonusly: GREEN
Driving clause (standard, favorable): "Carve-outs: information that (a) is or becomes publicly available through no breach, (b) was known prior to disclosure, (c) is received from a third party without duty of confidence, (d) is independently developed, or (e) must be disclosed by law or court order."
Reasoning: Mutual, full standard carve-outs, DE law, 2yr term + 3yr survival, no non-solicit/non-compete/exclusivity — clean standard approval.

NDA-2 — [PARTY B] and Bonusly: YELLOW
Driving clause (embedded non-solicit): "During the term of this Agreement and for eighteen (18) months thereafter, neither party shall solicit for employment or hire any employee of the other party with whom it came into contact under this Agreement."
Reasoning: Otherwise standard (mutual, full carve-outs, DE law, 3yr term), but the 18-month post-term employee non-solicit restricts Bonusly hiring and needs counsel review of scope ("came into contact") and duration.

NDA-3 — [PARTY C] and Bonusly: RED
Driving clause (non-compete): "For a period of three (3) years following the Effective Date, Recipient shall not, directly or indirectly, engage in or provide services to any business that competes with Discloser's business."
Reasoning: One-way (favoring [PARTY C] only) + 3-year non-compete on Bonusly's core business + "No confidentiality carve-outs are stated in this Agreement" + foreign governing law ("the laws of the Republic of Ireland ... exclusive jurisdiction of its courts") — four independent red flags requiring full legal review.

Notes
- No arithmetic required; all terms quoted verbatim from ndas.md.
- Per Bonusly deal desk policy, NDAs are transactional, but NDA-2 and NDA-3 are non-standard and must be flagged to Ninad before signature; nothing goes externally without human review.
- Sound too good to be true? Ask Deal Desk to verify or consult with our external counsel.