← back to all models ← RevenueBench home
nex-agi

nex-agi/nex-n2-pro

Mean 0.929 · 10/14 perfect tests · $0.05 total · 14/40 tests completed — remaining tests discarded from study (endpoint timeouts after dedicated retries) · modalities in: text · out: text · each card: the prompt → the correct answer (gold) → this model's actual answer

What do the modalities mean?

Modalities in means what you can send this model: text, images, files, video, or audio. Modalities out means what it can send back. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what this model can handle in real work.

data-crm 1 tests
snowflake-sql-trapped-catalog0.83
Plain-English question: Write one Snowflake query for first-month usage per customer using only the provided, trap-laden catalog.
Prompt — what the model saw
Using ONLY the catalog provided, write one Snowflake SQL query returning, per customer company, for its first calendar month as a customer: unique givers, recognition count, successful redemption count. Read the catalog notes carefully: some tables are stale or unpopulated and must be avoided, and one documented filter (excluding deleted givers) must NOT be applied to historical giving counts. Put a comment above each FROM or JOIN explaining the table choice. Return only fenced SQL.
cost $0.0017 · 5075s · in 1,067 / out 1,018 tok
✓ avoids_stale_deals✓ avoids_unpopulated_live✓ uses_cohort_summary✗ succeeded_filter✓ no_giver_deletion_filter✓ returns_sql

✓ Correct Answer (ground truth)

Plain-English answer: A single correct SQL query that uses the right catalog tables and avoids the trap columns — graded on the exact tables and columns.
The answer in one line: use_table = PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
avoid_stalePRODUCTION.HUBSPOT.DEALS (stale)
avoid_unpopulatedV2_LIVE.OBJECTS_DEALS (unpopulated)
forbidden_filterdeleted-giver exclusion must NOT be applied
outputone fenced SELECT with per-FROM/JOIN table-choice comment
required_filtersucceeded redemptions only
use_tablePRODUCTION.PLG.COMPANY_COHORT_SUMMARY
graded checksavoids_stale_deals · avoids_unpopulated_live · uses_cohort_summary · succeeded_filter · no_giver_deletion_filter · returns_sql

Model's actual answer

show the model's full answer
```sql
WITH customer_first_month AS (
    SELECT
        CUSTOMER_ID AS customer_company,
        DATE_TRUNC('MONTH', TO_DATE(FIRST_SUB_PAYMENT_DATE)) AS first_calendar_month_as_customer
    FROM PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
    -- FROM COMPANY_COHORT_SUMMARY: documented self-serve company source with first payment and M1 recognition metrics.
    WHERE FIRST_SUB_PAYMENT_DATE IS NOT NULL
),

giving AS (
    SELECT
        customer_company,
        COUNT(DISTINCT giver_id) AS unique_givers,
        COUNT(*) AS recognition_count
    FROM PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
    -- FROM COMPANY_COHORT_SUMMARY: documented source for M1_GIVES; no deleted-giver exclusion is applied because catalog says that filter must not be used for historical giving counts.
    GROUP BY 1
),

redemptions AS (
    SELECT
        customer_company,
        COUNT(*) AS successful_redemption_count
    FROM PRODUCTION.PLG.COMPANY_COHORT_SUMMARY
    -- FROM COMPANY_COHORT_SUMMARY: documented source for M1_REDEMPTIONS; PRODUCTION.DEPRECATED_RECOGNITION.REDEMPTION_RECORDS_V2 is documented as requiring confirmation for finance-grade counts, so this query uses the cataloged cohort summary.
    GROUP BY 1
)

SELECT
    c.customer_company,
    c.first_calendar_month_as_customer,
    COALESCE(g.unique_givers, 0) AS unique_givers,
    COALESCE(g.recognition_count, 0) AS recognition_count,
    COALESCE(r.successful_redemption_count, 0) AS successful_redemption_count
FROM customer_first_month c
-- FROM customer_first_month: derives each company's first calendar month as a customer from FIRST_SUB_PAYMENT_DATE.
LEFT JOIN giving g
    -- JOIN giving: attaches first-month giving counts without excluding deleted givers.
    ON g.customer_company = c.customer_company
LEFT JOIN redemptions r
    -- JOIN redemptions: attaches first-month redemption counts.
    ON r.customer_company = c.customer_company
ORDER BY
    c.customer_company;
```
deal-intelligence 2 tests
next-to-close0.75
Plain-English question: Which 3 deals are closest to signature right now?
Prompt — what the model saw
Which 3 deals are closest to signature right now? Use the CRM data and the Slack thread; Slack often has the real status before the CRM does. For each: deal, amount, why it is close, what is left. Under 250 words.
cost $0.0024 · 5018s · in 1,607 / out 1,628 tok
✓ exactly_three✗ top3_named✓ slack_first✓ blocked_excluded

✓ Correct Answer (ground truth)

Plain-English answer: Deal-547B2B ($11,200), Deal-B7EBD1 ($9,000), and Deal-403845 ($9,000) — all DS5 COMMIT deals closing September 10-11, with what is left on each.
blocked_excludeddeal_alias=Deal-2465CE; reason=champion left + procurement freeze (Slack #sales-sparring 2026-09-04)
exactly_three1
ranking_ruleSlack signature-imminent deal ranks first; remaining ranked by stage (DS5>DS4), forecast (COMMIT>BEST_CASE>PIPELINE), close_date asc, amount desc
top3rank=1; deal_alias=Deal-547B2B; amount=11200; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11 | rank=2; deal_alias=Deal-B7EBD1; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-10 | rank=3; deal_alias=Deal-403845; amount=9000; stage=DS5; forecast_category=COMMIT; close_date=2026-09-11
graded checksexactly_three · top3_named · slack_first · blocked_excluded

Model's actual answer

show the model's full answer
Closest to signature:

1. Deal-547B2B — $11,200  
Why close: Slack says redlines came back clean, signing page is out, and VP People said they are signing tomorrow. Arithmetic: thread date 2026-09-04; “tomorrow” = 1 day. CRM DS5/COMMIT close date 2026-09-11 = 7 days.  
Left: signature / contract execution.

2. Deal-403845 — $9,000  
Why close: Slack says the order form is with finance; CRM is DS5/COMMIT. Arithmetic: 2026-09-11 - 2026-09-04 = 7 days.  
Left: finance approval and signature.

3. Deal-A2B47C — $6,360  
Why close: Slack says it is still warm with normal legal-review pace; CRM is DS5/COMMIT. Arithmetic: 2026-09-11 - 2026-09-04 = 7 days.  
Left: complete legal review and signature.

Excluded: Deal-2465CE — CRM shows $5,400 and 2026-09-10, but Slack says the champion left, procurement froze new vendors, and it is blocked until they re-staff / Q4.
product-gap-detectionPASS 1.00
Plain-English question: Which prospect complaints are true product gaps versus rollout or enablement issues?
Prompt — what the model saw
Surface true product gaps raised by prospects in active deals. For each candidate: quote the prospect line, classify as real gap, plan gate, or rollout/enablement issue using the product docs excerpt, and name the deal and amount. Only real gaps go in the summary. Rep language does not count as prospect voice.
cost $0.0018 · 3961s · in 796 / out 1,141 tok
✓ real_gaps✓ plan_gate_classified✓ enablement_classified✓ gap_summary

✓ Correct Answer (ground truth)

Plain-English answer: Two real gaps (SCIM provisioning, ADP Workforce Now integration) separated from plan gates and enablement issues, each with the prospect quote and deal.
allowed_newTG-001, TG-002, TG-003, TG-004, TG-005, SCIM user provisioning, ADP Workforce Now integration, Custom report builder
enablement_issuestranscript_id=TG-004; deal_alias=Deal-84DBA6; amount=16000; feature=Slack integration recognition adoption; quote=We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.; why_enablement=Slack integration exists on all plans; managers were never trained
excluded_rep_voicetranscript_id=TG-005; line=Full transparency — we don't have a native mobile app for hourly workers yet, though it's on the roadmap.; reason=rep voice, not prospect-raised
plan_gatestranscript_id=TG-003; deal_alias=Deal-CFE7F4; amount=32000; feature=Custom report builder; tier_required=Enterprise; quote=I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?
real_gapstranscript_id=TG-001; deal_alias=Deal-EC3025; amount=62000; feature=SCIM user provisioning; quote=We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.; why_real=not listed in product docs at any tier | transcript_id=TG-002; deal_alias=Deal-D0D6B5; amount=23800; feature=ADP Workforce Now integration; quote=Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.; why_real=supported HRIS list is Workday/BambooHR/Gusto/Rippling only
graded checksreal_gaps · plan_gate_classified · enablement_classified · gap_summary

Model's actual answer

show the model's full answer
Candidate review:

| Transcript | Deal | Amount | Prospect line | Classification | Evidence from product docs |
|---|---:|---:|---|---|---|
| TG-001 | Deal-EC3025 | Not provided | “We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.” | Real gap | Docs state: “SCIM user provisioning … are NOT currently listed as supported capabilities.” |
| TG-002 | Deal-D0D6B5 | Not provided | “Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.” | Real gap | Docs state: “ADP Workforce Now integrations are NOT currently listed as supported capabilities.” |
| TG-003 | Deal-CFE7F4 | Not provided | “I tried to build a custom report combining recognition frequency with turnover and couldn't — is that possible on our current plan?” | Plan gate | Docs show “Custom report builder” is Enterprise only. |
| TG-004 | Deal-84DBA6 | Not provided | “We turned on the Slack integration in July, but half our managers still haven't given a single recognition — nobody ever trained them on it.” | Rollout/enablement issue | Docs show “Slack and Microsoft Teams integration” is available on all plans. |
| TG-005 | Deal-36C33F | Not provided | “Good to know. The web version should be fine for our office staff for now.” | Not a prospect-raised gap | Prospect accepted the web version; no gap raised by prospect. Rep-only mobile-app statement is excluded. |

Arithmetic:
- Prospect-raised candidate lines reviewed: 4
- Real product gaps: 2
- Plan gates: 1
- Rollout/enablement issues: 1
- Prospect lines not treated as gaps: 1
- Real-gap deal amount total: not calculable because amount is not provided for any deal.

Summary of true product gaps:

1. Deal-EC3025 — SCIM user provisioning  
   Prospect quote: “We need SCIM user provisioning — our security team will not approve any tool that requires manual account management.”  
   Amount: Not provided.

2. Deal-D0D6B5 — ADP Workforce Now integration  
   Prospect quote: “Our payroll runs on ADP Workforce Now and I don't see ADP anywhere in your integration list — that's a dealbreaker for us.”  
   Amount: Not provided.
customer-success 2 tests
expansion-kitPASS 1.00
Plain-English question: Build the expansion kit for the customer who just replied to our outreach.
Prompt — what the model saw
A customer replied to outreach (reply included). Build the expansion kit: seat coverage (licensed over headcount), two lines on usage health, headroom in seats and ARR at the current per-seat rate, who replied and whether they can buy, the right buyer from our contacts, and a reply email under 150 words that moves toward the right person without being pushy and cites one real usage fact.
cost $0.0022 · 1031s · in 519 / out 1,696 tok
✓ coverage_pct✓ headroom_arr✓ right_buyer✓ replier_not_buyer✓ usage_fact✓ email_under_150

✓ Correct Answer (ground truth)

Plain-English answer: 37.5% seat coverage, 250 headroom seats worth $15,000 ARR at $60 a seat, usage up 43.2% in six months — and the replier cannot buy; Dana R. (VP People) is the right buyer.
The answer in one line: headroom_seats = 250; coverage_pct = 37.5; company_alias = C-0DDFC9A7 (+3 more below)
headroom_seats250
coverage_pct37.5
company_aliasC-0DDFC9A7
headroom_arr15000.00
per_seat_annual_rate60.00
replierMaria S.
replier_can_buy0
right_buyername=Dana R.; title=VP People
usage_factsactive_users_2026_08=126; active_users_2026_03=88; growth_pct_6m=43.2; licensed_seat_utilization_pct=84
usage_health_linesMonthly active users grew from 88 (2026-03) to 126 (2026-08), +43.2% over 6 months., Last month 126 of 150 licensed seats were active (84.0% utilization), approaching the seat cap.
graded checkscoverage_pct · headroom_arr · right_buyer · replier_not_buyer · usage_fact · email_under_150

Model's actual answer

show the model's full answer
Company alias: C-0DDFC9A7  
Deal alias: not provided

Seat coverage
- Licensed seats / headcount = 150 / 400 = 37.5%
- Headroom in seats = 400 - 150 = 250 seats
- Current per-seat rate = $9,000 / 150 = $60/seat
- ARR headroom = 250 seats × $60/seat = $15,000

Usage health
- MAUs grew every month: 88 → 95 → 102 → 110 → 118 → 126; Mar-to-Aug increase = 126 - 88 = +38 users.
- Aug usage = 126 / 150 licensed seats = 84.0% licensed-seat utilization.

Who replied and whether they can buy
- Maria S., People Operations Coordinator, replied on 2026-09-02.
- She said she is not the right person for purchasing decisions, so she cannot buy.

Right buyer from our contacts
- Dana R., VP People, last engaged 2026-05-18.
- Maria said budget and seat expansion sit with Dana R.

Reply email under 150 words
Hi Maria,

Thanks for sharing — glad to hear the team loves Bonusly. Your usage has grown from 88 MAUs in Mar to 126 in Aug, so it may be helpful to compare that trend with Dana’s seat-expansion planning.

If you’re comfortable, could you introduce me to Dana? I’ll keep it brief and share the usage snapshot she’s been asking about.

Best,  
Cole
onboarding-statusPASS 1.00
Plain-English question: What is done and what is missing in this customer's onboarding before the mid-onboarding call?
Prompt — what the model saw
Prep the CSM for the mid-onboarding call: what is complete (integrations connected, allowance set, admins added, first recognition given, first redemption), what is not, early engagement signals, and the three things to cover on the call. Mark nothing complete without a data field that shows it.
cost $0.0028 · 63s · in 531 / out 2,209 tok
✓ complete_items✓ incomplete_items✓ three_items✓ grounded

✓ Correct Answer (ground truth)

Plain-English answer: 25 days live: Slack connected, allowance set, admins added — but the HRIS integration is not connected, with early engagement signals and three call priorities.
The answer in one line: days_live = 25; company_alias = C-0D284E42
days_live25
call_itemsConnect the HRIS integration (integration_hris is empty) to automate user provisioning and offboarding., Get to first redemption (first_redemption_at is empty): walk admins through the rewards catalog and confirm allowance funding., Build the giving habit: daily active givers averaged 4.3 in week 1 vs 13.0 last week - agree on manager nudges and a recognition cadence.
checklist{"slack_integration_connected": {"complete": true, "field": "integration_slack", "value": "2026-08-12"}, "hris_integration_connected": {"complete": false, "field": "integration_hris", "value": ""}, "a
company_aliasC-0D284E42
engagement_signalsavg_active_givers_week1=4.3; avg_active_givers_last_week=13; latest_active_givers=15; days_with_activity=25
graded checkscomplete_items · incomplete_items · three_items · grounded

Model's actual answer

show the model's full answer
Company: C-0D284E42

## Complete

- Slack integration connected: `integration_slack = 2026-08-12`
- Allowance set: `allowance_set = 2026-08-13`
- Admins added: `admins_added = 2`
- First recognition given: `first_recognition_at = 2026-08-15 14:22`

## Not complete / missing

- HRIS integration not connected: `integration_hris` is blank.
- First redemption not complete: `first_redemption_at` is blank.

## Early engagement signals

Usage data covers 25 days: 2026-08-11 through 2026-09-04.

- Total active-giver-days: `214`
- Average active givers/day: `214 / 25 = 8.56`
- First 7 days: `30 / 7 = 4.29` active givers/day
- Last 7 days: `91 / 7 = 13.00` active givers/day
- First-to-latest active givers: `15 - 3 = +12`
- Growth multiple: `15 / 3 = 5.0x`
- Peak active givers: `15` on `2026-09-03` and `2026-09-04`

## Three things to cover on the call

1. HRIS integration is still missing; confirm owner and next step for `integration_hris`.
2. First redemption is still missing despite recognition happening on `2026-08-15 14:22`; identify why no `first_redemption_at` exists yet.
3. Reinforce the active-giver growth trend: active givers rose from `3` to `15`, with the last 7-day average at `13.00/day` versus the first 7-day average at `4.29/day`.
marketing 3 tests
case-study-matchPASS 1.00
Plain-English question: Which three existing customers are the best case-study matches for this prospect?
Prompt — what the model saw
Given the prospect, rank the three most similar existing customers with public case studies for use as social proof. Similarity on industry, size band, use case, region. Cite the fields that drove each match. Do not name any customer without a case study and never include billing data or contact names.
cost $0.0027 · 356s · in 720 / out 2,114 tok
✓ top3✓ decoys_excluded✓ cites_fields✓ no_billing

✓ Correct Answer (ground truth)

Plain-English answer: Three ranked matches on industry, size, use case, and region (e.g. C-11C31562), with the higher-scoring decoy explicitly excluded for a mismatched field.
The answer in one line: prospect_alias = C-82AF3719
constraintsno_billing_data, no_contact_names, only_customers_with_has_case_study_true_named
decoys_excludedcompany_alias=C-D6217CAA; similarity_score=4; reason=no_case_study | company_alias=C-C153868F; similarity_score=4; reason=no_case_study
prospect_aliasC-82AF3719
prospect_fields_usedindustry, size_band, use_case, region
top3company_alias=C-11C31562; similarity_score=3 | company_alias=C-64171065; similarity_score=3 | company_alias=C-A13C193D; similarity_score=3
graded checkstop3 · decoys_excluded · cites_fields · no_billing

Model's actual answer

show the model's full answer
Filtered to customers with `has_case_study=true`.

Scoring: 1 point per exact match on `industry`, `size_band`, `use_case`, `region`; max = 4/4.  
Prospect `C-82AF3719`: Technology / Mid-Market / employee_recognition / NA-West.

1. `C-64171065` — 3/4  
   - Matched fields: `industry=Technology`, `size_band=Mid-Market`, `use_case=employee_recognition`  
   - Mismatch: `region=NA-East` vs `NA-West`  
   - Arithmetic: 1 + 1 + 1 + 0 = 3/4

2. `C-A13C193D` — 3/4  
   - Matched fields: `industry=Technology`, `size_band=Mid-Market`, `region=NA-West`  
   - Mismatch: `use_case=retention` vs `employee_recognition`  
   - Arithmetic: 1 + 1 + 0 + 1 = 3/4

3. `C-11C31562` — 3/4  
   - Matched fields: `size_band=Mid-Market`, `use_case=employee_recognition`, `region=NA-West`  
   - Mismatch: `industry=Manufacturing` vs `Technology`  
   - Arithmetic: 0 + 1 + 1 + 1 = 3/4

No public-case-study customer matched all four fields.
channel-attribution-mathPASS 1.00
Plain-English question: Compute cost per SQM, cost per SQO, and pipeline per dollar for every channel over six months.
Prompt — what the model saw
Compute channel performance for the trailing 6 months. Paid: spend, SQMs, SQOs, cost per SQM, cost per SQO, SQM-to-SQO rate, pipeline amount, pipeline per dollar. Organic: volume, SQO rate, pipeline. Report a channel with spend and zero SQMs as undefined, not zero. Flag rows where SQO date precedes SQM date. Finish with a reallocation recommendation and your confidence given sample sizes.
cost $0.0086 · 5961s · in 147 / out 7,405 tok
✓ undefined_channel✓ undefined_not_zero✓ flagged_rows✓ best_channel✓ reallocation

✓ Correct Answer (ground truth)

Plain-English answer: LinkedIn Ads at $960/SQM and $3,000/SQO, organic search at $90,000 pipeline on zero spend, paid social flagged at $18,000 with zero SQMs, and a channel winner named.
The answer in one line: best_pipeline_per_dollar_channel = paid_search
best_pipeline_per_dollar_channelpaid_search
confidence_notewebinars and referral have small SQM samples (<20); treat rates as directional.
flagged_rows_sqo_before_sqmCT-000041, CT-000044
months2026-03, 2026-04, 2026-05, 2026-06, 2026-07, 2026-08
per_channel{"linkedin_ads": {"spend_usd": 24000, "sqms": 25, "sqos": 8, "pipeline_amount": 96000, "cost_per_sqm": 960.0, "cost_per_sqo": 3000.0, "sqm_to_sqo_rate": 0.32, "pipeline_per_dollar": 4.0}, "organic_sea
reallocation_recommendationfrom=paid_social; to=paid_search; rationale=paid_social has spend but zero SQMs (cost per SQM undefined, not zero); paid_search has the highest pipeline per dollar.
undefined_channelspaid_social
graded checksundefined_channel · undefined_not_zero · flagged_rows · best_channel · reallocation

Model's actual answer

show the model's full answer
Trailing 6 months: 2026-03 through 2026-08.

### Paid channels

| Channel | Spend | SQMs | SQOs | Cost/SQM | Cost/SQO | SQM→SQO rate | Pipeline | Pipeline/$ |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| paid_search | $36,000 | 40 | 18 | $900 | $2,000 | 45.0% | $720,000 | $20.00 |
| linkedin_ads | $24,000 | 25 | 8 | $960 | $3,000 | 32.0% | $96,000 | $4.00 |
| paid_social | $18,000 | 0 | 0 | undefined | undefined | undefined | $0 | $0.00 |
| webinars | $9,000 | 12 | 5 | $750 | $1,800 | 41.7% | $60,000 | $6.67 |

### Organic channels

| Channel | Volume / SQMs | SQOs | SQO rate | Pipeline |
|---|---:|---:|---:|---:|
| organic_search | 30 | 10 | 33.3% | $90,000 |
| referral | 15 | 6 | 40.0% | $48,000 |

### SQO date precedes SQM date

Flagged rows:
- CT-000044, linkedin_ads: SQM date = 2026-07-23; SQO date = 2026-07-18; pipeline_amount = $12,000.
- CT-000041, linkedin_ads: SQM date = 2026-06-14; SQO date = 2026-06-09; pipeline_amount = $12,000.

### Reallocation recommendation

Shift paid budget toward paid_search and away from paid_social and linkedin_ads.

Reason:
- paid_search has the best pipeline/$ at $20.00, strongest SQM→SQO rate at 45.0%, and $720,000 pipeline on $36,000 spend.
- webinars has $6.67 pipeline/$ and 41.7% SQM→SQO rate, but lower scale.
- linkedin_ads has $4.00 pipeline/$, 32.0% SQM→SQO rate, and 2 data-quality flags.
- paid_social has spend of $18,000 and zero SQMs, so cost/SQM, cost/SQO, and SQM→SQO rate are undefined.

Confidence: Medium-low. Sample sizes are small: paid_search has 40 SQMs, linkedin_ads has 25 SQMs, paid_social has 0 SQMs, webinars has 12 SQMs, organic_search has 30 SQMs, and referral has 15 SQMs. The recommendation is directionally clear from the provided data, but paid_social and webinars are especially noisy due to low/no SQM counts.
competitor-battlecard-updatePASS 1.00
Plain-English question: Update the Rivally battlecard with sourced pricing and where each side wins.
Prompt — what the model saw
Update the battlecard for the competitor. Sections: one-line positioning, pricing with source and date (newer source wins, note the conflict), where they win, where we win, objections and responses, recent changes, our 12-month win/loss record against them. Cite a snippet id for every factual claim. Rep opinion on a call is not a fact about the competitor. Anything from the old card you cannot re-source gets marked unverified.
cost $0.0144 · 1111s · in 1,572 / out 13,573 tok
✓ competitor_named✓ pricing✓ win_loss✓ unverified_marked✓ citations

✓ Correct Answer (ground truth)

Plain-English answer: Rivally at $7 per user per month (the August source wins over the stale $5, conflict noted), a 13-7 win/loss record, and two unverified old claims flagged.
The answer in one line: competitor = Rivally
citation_ids_available25 items — e.g. S01, S02, S03, S04
competitorRivally
pricingvalue_per_user_per_month=7; source_snippet=S17; source_date=2026-08-12
rep_opinion_snippet_ids_not_factsS09, S21
unverified_old_claimsRivally lacks a Slack integration., Rivally was acquired by WorkHuman in 2025.
win_loss_12mowins=13; losses=7
graded checkscompetitor_named · pricing · win_loss · unverified_marked · citations

Model's actual answer

show the model's full answer
# Battlecard: Rivally

## One-line positioning
Rivally is a recognition platform centered on a points-based recognition feed, with mid-market setup evidence and EU expansion/data-residency messaging (S02, S04, S05, S11, S12, S15).

## Pricing — newer source wins; note conflicts
- Current public pricing to use: Recognition Starter is $7/user/month, annual billing required (S17, 2026-08-12 pricing_page).
- Older pricing conflict: $5/user/month, annual billing required (S03, 2026-01-20 pricing_page; S08, 2026-04-01 pricing_page).
- Deal-note pricing conflict: $6.50/user/month annual quote for a 500-seat prospect (S13, 2026-06-02 call_notes); $7/user/month list with 15% discount for a 3-year term (S18, 2026-08-14 call_notes).

## Where Rivally wins
- Recognition engagement: points-based recognition feed praised (S02); recognition feed described as engaging (S16).
- Implementation/integration: mid-market setup took under a week, and Slack integration worked out of the box (S04).
- EU coverage: Rivally pitched EU data residency (S05); EU enterprise reviewer praised distributed EU teams and multi-language support (S12); Dublin office opened and EU data residency became generally available (S15).
- Support: support response time praised as under 4 hours (S22).

## Where we win / supported win themes
- Analytics depth/exportability: an 800-seat prospect picked Bonusly over Rivally citing analytics depth (S25); Rivally analytics are cited as limited (S02), basic compared to enterprise tools (S07), and CSV-only for exports (S20).
- Enterprise admin/security gaps to exploit: Rivally lacks SCIM provisioning and manual user management is painful (S10); admin tooling lags peers (S16); admin console lacks bulk recognition editing (S24).
- No other specific Bonusly win reasons are provided in S01-S25.

## Objections and responses
| Objection | Evidence | Response |
|---|---|---|
| Rivally’s recognition feed is engaging. | Points-based feed praised (S02); feed described as engaging (S16). | Acknowledge, then qualify analytics/reporting depth and export needs: limited analytics (S02), basic dashboards vs enterprise tools (S07), CSV-only analytics exports (S20). |
| Rivally is easy to implement and Slack works. | Setup under a week; Slack worked out of the box (S04). | Acknowledge, then probe admin/security requirements: no SCIM and painful manual user management (S10), admin tooling lags peers (S16), no bulk recognition editing (S24). |
| Rivally supports EU teams and data residency. | EU data residency pitched (S05); EU reviewer praised distributed EU teams and multi-language support (S12); EU data residency generally available (S15). | Acknowledge, then test EMEA rewards catalog breadth: EMEA catalog is thinner than US catalog (S14). |
| Rivally is cheaper. | $7/user/month Starter with annual billing required (S17); $7/user/month list with 15% discount for 3-year term (S18). | Normalize terms: annual billing required (S17), 3-year discount term (S18), and Pulse is priced as an add-on, not bundled (S23). |
| Rivally has fast support. | Support response time praised as under 4 hours (S22). | Acknowledge, then pivot to analytics/admin gaps (S07, S10, S20, S24). |

## Recent changes
- Rivally raised a $40M Series C led by Northgate Ventures (S01, 2025-11-04).
- Rivally launched Rivally Pulse, a lightweight engagement survey add-on (S06, 2026-03-05).
- Rivally hired an ex-Workday VP EMEA to lead European expansion (S11, 2026-05-09).
- Rivally opened a Dublin office and announced EU data residency generally available (S15, 2026-07-01).
- Rivally updated Recognition Starter to $7/user/month with annual billing required (S17, 2026-08-12).
- Rivally announced Microsoft Teams app v2 in public preview (S19, 2026-08-20).
- Rivally Pulse exited beta and is priced as an add-on, not bundled (S23, 2026-09-01).
- Rivally admin console still lacks bulk recognition editing (S24, 2026-09-02).
- 800-seat prospect picked Bonusly over Rivally citing analytics depth (S25, 2026-09-03).

## Our 12-month win/loss record against Rivally
Source for all rows: deals_with_competitor.csv; this file has no snippet_id column.

| Month | Record | Deal aliases |
|---|---:|---|
| 2025-09 | 1-1 | Win: Deal-072E31; Loss: Deal-7767F5 |
| 2025-10 | 2-0 | Win: Deal-A9FD43, Deal-F65C8F |
| 2025-11 | 1-1 | Win: Deal-7AA785; Loss: Deal-D263E0 |
| 2025-12 | 1-1 | Win: Deal-44C524; Loss: Deal-935746 |
| 2026-01 | 2-0 | Win: Deal-0D0CD6, Deal-E46EAB |
| 2026-02 | 2-0 | Win: Deal-D5B790, Deal-1D2392 |
| 2026-03 | 1-1 | Win: Deal-5C636E; Loss: Deal-9066A6 |
| 2026-04 | 0-2 | Loss: Deal-5645A5, Deal-72A02F |
| 2026-05 | 0-1 | Loss: Deal-C6FFAA |
| 2026-06 | 1-0 | Win: Deal-67BE14 |
| 2026-07 | 1-0 | Win: Deal-1B6969 |
| 2026-08 | 1-0 | Win: Deal-F03E7B |

Arithmetic: 13 wins + 7 losses = 20 total deals. Win rate = 13 / 20 = 65%. Loss rate = 7 / 20 = 35%.

## Old-card status
- Re-sourced: points-based recognition feed and mid-market setup evidence (S02, S04).
- Re-sourced but stale: $5/user/month annual billing (S03, S08); newer source shows $7/user/month annual billing (S17).
- Contradicted: “Rivally lacks a Slack integration” — S04 says Slack integration worked out of the box.
- Re-sourced: strong for distributed EU teams and multi-language support (S12); EU data residency generally available (S15).
- Unverified: “Rivally was acquired by WorkHuman in 2025” — no supporting snippet in S01-S25.

Rep-opinion exclusions: S09 and S21 are not used as competitor facts; both are call-note opinions rather than verified competitor facts (S09, S21).
communication 4 tests
stalled-deal-followup-emailPASS 1.00
Plain-English question: Write a follow-up email for this stalled deal — under 120 words, one prior touch, one ask.
Prompt — what the model saw
Write a follow-up email for this stalled deal. Under 120 words. Reference one specific prior touch from the engagement history, make one clear ask, and match the tone of the prior thread. No invented details.
cost $0.0011 · 2996s · in 437 / out 559 tok
✓ word_cap✓ references_touch✓ single_ask

✓ Correct Answer (ground truth)

Plain-English answer: An email that references the July 28 milestone-awards demo, makes one clear ask, and invents nothing — Deal-0D2F7A, 31 days since last engagement.
The answer in one line: amount = 5100; word_cap = 120; days_since_last_engagement = 31; deal_alias = Deal-0D2F7A (+4 more below)
amount5,100
days_since_last_engagement31
word_cap120
allowed_facts14 items — e.g. Deal-0D2F7A, $5,100, DS3, Alex Franklin
allowed_newENG-00001, ENG-00002, ENG-00003, ENG-00004, ENG-00005
deal_aliasDeal-0D2F7A
must_referenceengagement_id=ENG-00002; date=2026-07-28; type=meeting; why=most substantive prior touch; concrete prospect reactions to milestone awards and points catalog
ownerAlex Franklin
single_cta1
stageDS3
tonewarm, concise, professional (matches prior thread)
graded checksword_cap · references_touch · single_ask

Model's actual answer

show the model's full answer
Subject: Deal-0D2F7A follow-up

Hi,

Following up on Deal-0D2F7A after the Aug. 5 email recapping the July 28 demo and pricing for 150 seats.

Given the People team’s interest in automated milestone awards and the points catalog, would it be useful to revisit pricing?

Best,  
Alex
funnel-digest-paragraphPASS 1.00
Plain-English question: Write the weekly GTM digest — one paragraph per team, every number traceable.
Prompt — what the model saw
Write the weekly GTM digest: one paragraph per team (Marketing, Sales, CS, PLG) in Bonusly brand voice. Every number must trace to the extracts. No customer contact names, no employee performance commentary, no compensation. Four paragraphs, under 400 words total.
cost $0.0023 · 1435s · in 266 / out 1,821 tok
✓ four_paragraphs✓ word_cap✓ teams

✓ Correct Answer (ground truth)

Plain-English answer: Four paragraphs (Marketing, Sales, CS, PLG) using only the allowed numbers, no contact names, no performance commentary, under 400 words.
The answer in one line: word_cap = 400
word_cap400
allowed_numbers16 items — e.g. 2, 3, 9, 08
bannedcustomer contact names, employee performance commentary, compensation
required_paragraphsMarketing, Sales, CS, PLG
graded checksfour_paragraphs · word_cap · teams

Model's actual answer

show the model's full answer
Marketing: SQM attainment came in at 46 / 52 = 88.5%, with the webinar channel driving 18 / 46 = 39.1% of SQMs. Paid search held at $150 cost per SQM, keeping spend efficiency steady.

Sales: Sales converted 14 SQOs and set 9 DS2 meetings. New pipeline created was $310,000, and same-quarter close count was 3. No SQO conversion-rate denominator was provided, so no rate is calculated.

CS: Renewals saved this week were 2, and Team NPS moved to 61. There are 3 open red-flag accounts heading into next week, so the focus is clear: protect the base and remove friction fast.

PLG: PLG added 412 new signups, with activation at 31%. 38 companies reached the aha moment of 10 recognition gives. No overlap between signups and companies was provided, so no additional conversion rate is calculated.
partner-digest-section0.67
Plain-English question: Write the partner digest so every partner in the roster appears exactly once.
Prompt — what the model saw
Write the partner digest for the period. Every partner in the roster appears once: active plays get status with numbers, quiet partners get one line. Partner-sourced pipeline count and amount must match the deal data. Nothing omitted, nothing added.
cost $0.0019 · 24s · in 551 / out 1,369 tok
✓ all_partners✗ pipeline_amount✓ quiet_one_line

✓ Correct Answer (ground truth)

Plain-English answer: Four active partners with status and numbers, four quiet partners with one line each, and partner-sourced pipeline that reconciles to the deal data.
The answer in one line: rule = every roster partner named exactly once; quiet partners get one line; no partners outside roster.
active_no_dealsWorkWell Group
partner_sourced_pipelinecount=5; amount=550000
per_partner_pipeline{"Apex Rewards Co": {"count": 2, "amount": 275000}, "HRCloud Partners": {"count": 1, "amount": 140000}, "CultureBridge": {"count": 2, "amount": 135000}}
quiet_partnersRecogniq, KudosWave, PeopleFirst Advisors, TotalPerk
rosterpartner_name=Apex Rewards Co; status=active | partner_name=HRCloud Partners; status=active | partner_name=CultureBridge; status=active | partner_name=WorkWell Group; status=active | partner_name=Recogniq; status=quiet | partner_name=KudosWave; status=quiet … (+2 more)
ruleevery roster partner named exactly once; quiet partners get one line; no partners outside roster.
graded checksall_partners · pipeline_amount · quiet_one_line

Model's actual answer

show the model's full answer
Partner digest for 2026-08-24 to 2026-09-04

- Apex Rewards Co (active): 2 partner-sourced deals, $275,000. Arithmetic: Deal-DDAAF2 $180,000 + Deal-2CF33E $95,000 = $275,000. Status: co-webinar locked for 09-15; both sourced opps in DS1 and logged with UTM Source = Partner.

- HRCloud Partners (active): 1 partner-sourced deal, $140,000. Arithmetic: Deal-F1CDA5 $140,000 = $140,000. Status: integration referral closed security review; opp moved to DS2.

- CultureBridge (active): 2 partner-sourced deals, $135,000. Arithmetic: Deal-096E1D $60,000 + Deal-067213 $75,000 = $135,000. Status: lunch-and-learn produced two sourced opps; both early stage.

- WorkWell Group (active): 0 partner-sourced deals, $0. Arithmetic: no partner_deals.csv rows for WorkWell Group. Status: wants to restart the joint playbook in Q4; no sourced deals this period; planning call booked for 09-09.

- Recogniq (quiet): No sourced deals in partner_deals.csv; no Slack update in the period.

- KudosWave (quiet): No sourced deals in partner_deals.csv; no Slack update in the period.

- PeopleFirst Advisors (quiet): No sourced deals in partner_deals.csv; no Slack update in the period.

- TotalPerk (quiet): No sourced deals in partner_deals.csv; no Slack update in the period.
rfp-answer-from-policyPASS 1.00
Plain-English question: Answer the 10 RFP items from the policy excerpts only — no guessing.
Prompt — what the model saw
Answer the 10 questionnaire items from the policy excerpts only. Cite the excerpt id for every answer. Where the excerpts do not answer the question, mark it unanswerable and say what document would; do not guess. Tone: vendor response, direct.
cost $0.0014 · 1420s · in 522 / out 877 tok
✓ answered_count✓ unanswerable_flagged✓ citations✓ no_guess

✓ Correct Answer (ground truth)

Plain-English answer: Seven answered with excerpt citations (Q1-Q7), three flagged unanswerable with the document that would answer them (Q8-Q10).
The answer in one line: answered_count = 7; rule = unanswerable items must be flagged and the document that would answer them named; do not guess.
answered_count7
answer_map{"Q1": ["P1"], "Q2": ["P1"], "Q3": ["P2"], "Q4": ["P3"], "Q5": ["P4"], "Q6": ["P5"], "Q7": ["P6"]}
ruleunanswerable items must be flagged and the document that would answer them named; do not guess.
unanswerableQ8, Q9, Q10
valid_citation_idsP1, P2, P3, P4, P5, P6, P7, P8
graded checksanswered_count · unanswerable_flagged · citations · no_guess

Model's actual answer

show the model's full answer
| Item | Vendor response |
|---|---|
| Q1 | Yes. Customer data is encrypted at rest using AES-256. Cite: [P1] |
| Q2 | Yes. Data in transit is encrypted with TLS 1.2 or higher. Cite: [P1] |
| Q3 | Yes. Bonusly maintains a SOC 2 Type II report, renewed annually. It is available to prospects under NDA. Cite: [P2] |
| Q4 | Customer data is retained for the duration of the contract plus 90 days, then deleted from production systems. Cite: [P3] |
| Q5 | The current subprocessor list is published at trust.bonusly.com. Customers receive 30 days notice before a new subprocessor is added. Cite: [P4] |
| Q6 | Yes. SAML 2.0 single sign-on is available on the Enterprise plan. Cite: [P5] |
| Q7 | Production databases are backed up daily. Backups are encrypted and retained for 35 days. Cite: [P6] |
| Q8 | Unanswerable from the provided excerpts. No excerpt in [P1]–[P8] addresses ISO 27001 certification. This would be answered by a security/compliance certifications document or ISO 27001 certificate status page. |
| Q9 | Unanswerable from the provided excerpts. No excerpt in [P1]–[P8] addresses contractual uptime SLA percentage. This would be answered by the MSA, SLA schedule, or enterprise terms document. |
| Q10 | Unanswerable from the provided excerpts. No excerpt in [P1]–[P8] addresses HIPAA or a Business Associate Agreement. This would be answered by the DPA/BAA, HIPAA compliance documentation, or legal terms document. |
ops-maintenance 2 tests
runbook-from-slack-thread0.75
Plain-English question: Turn the resolved incident thread into a runbook.
Prompt — what the model saw
Turn the resolved incident thread into a runbook: ordered steps with the exact command or action taken, who ran it, how success was verified, and rollback for each step that changed state. Every step must trace to a message in the thread. Add nothing the thread does not contain; where the thread is unclear, mark the step as needing confirmation.
cost $0.0047 · 2984s · in 61 / out 4,296 tok
✗ ordered_steps✓ rollback✓ needs_confirmation✓ sourced

✓ Correct Answer (ground truth)

Plain-English answer: Ordered steps from the thread, each with the exact command, who ran it, how success was verified, and rollback where state changed.
The answer in one line: rule = no steps beyond what the thread contains; every step traces to a message id.
ambiguous_steporder=4; msg=M05; mark=needs-confirmation
ruleno steps beyond what the thread contains; every step traces to a message id.
stepsorder=1; msg=M01; actor=Bryce Harmon; action=acknowledge PagerDuty alert, take IC; state_change=False; verified_by=M01 | order=2; msg=M02; actor=Farid Osman; action=check queue depth; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M02 (48,213 pending) | order=3; msg=M04; actor=Farid Osman; action=pause auto-enqueue via feature flag; command=bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'; state_change=True; rollback=bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'; verified_by=M07 (queue depth falling) | order=4; msg=M05; actor=Elena Sinclair; action=cleared the dead set; state_change=True; needs_confirmation=thread gives no command or method for clearing the dead set | order=5; msg=M06; actor=Bryce Harmon; action=scale reward-worker to 6 replicas; command=kubectl scale deployment/reward-worker --replicas=6; state_change=True; rollback=kubectl scale deployment/reward-worker --replicas=3; verified_by=M07/M08 | order=6; msg=M08; actor=Cole Ingram; action=verify queue drained and error rate baseline; command=bundle exec rake sidekiq:queue_depth; state_change=False; verified_by=M08 (returns 0) … (+2 more)
graded checksordered_steps · rollback · needs_confirmation · sourced

Model's actual answer

show the model's full answer
Runbook — incident-2026-08-29-reward-queue

| Step | Trace | Who ran it | Exact command or action taken | Success verification | Rollback if state changed |
|---:|---|---|---|---|---|
| 1 | [M01] | Bryce Harmon | Acknowledged PagerDuty alert; took IC. | Verification is the message itself: “Acknowledging, taking IC.” | Needs confirmation: rollback/handoff action not stated. |
| 2 | [M02] | Farid Osman | Ran `bundle exec rake sidekiq:queue_depth`. | Output: reward queue at `48,213` pending jobs. Normal is under `500`. | N/A — observation. |
| 3 | [M03] | Farid Osman | Reported dead set status. | Dead set had `112` jobs, all `Redis::TimeoutError` from around `13:58`. | N/A — observation. |
| 4 | [M04] | Farid Osman | Paused enqueue: `bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'`. | Needs confirmation: no direct verification stated for the disable. Later queue decline in [M07]/[M08] is indirect only. | Rollback stated in [M04]: `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`. |
| 5 | [M05] | Elena Sinclair | Cleared out the dead set in the console. | Needs confirmation: exact command and verification not stated. | Needs confirmation: rollback not stated. |
| 6 | [M06] | Bryce Harmon | Scaled workers up: `kubectl scale deployment/reward-worker --replicas=6` from `3` to `6`. | [M07]: queue depth down to `9,400` and falling about `1,200/min`. [M08]: queue depth returned `0`; Datadog error rate back to baseline. | Rollback stated in [M06]: `kubectl scale deployment/reward-worker --replicas=3`. |
| 7 | [M07] | Farid Osman | Reported queue depth status. | Queue depth down to `9,400` and falling about `1,200/min`. | N/A — observation. |
| 8 | [M08] | Cole Ingram | Verified queue depth with `bundle exec rake sidekiq:queue_depth`; checked Datadog error rate. | `bundle exec rake sidekiq:queue_depth` returned `0`. Datadog error rate back to baseline. | N/A — verification. |
| 9 | [M09] | Bryce Harmon | Re-enabled enqueue: `bin/rails runner 'FeatureFlag.enable(:auto_recognition_enqueue)'`. | Verified `40` new jobs processed cleanly in the next `3` minutes. | Rollback command appears in [M04]: `bin/rails runner 'FeatureFlag.disable(:auto_recognition_enqueue)'`; needs confirmation that this is intended rollback for this step. |
| 10 | [M10] | Bryce Harmon | Scaled workers back down: `kubectl scale deployment/reward-worker --replicas=3`. | Queue stable at `0`; incident resolved. | Rollback stated in [M06]: `kubectl scale deployment/reward-worker --replicas=6`. |
feature-flag-state-summaryPASS 1.00
Plain-English question: Summarize what each feature flag controls and who has it on.
Prompt — what the model saw
Summarize the state of the listed feature flags: what each controls per the code excerpt, which companies or segments have it on, targeting rules, and any flag with no code reference or no targets. Company counts must match the export. Do not describe a flag that is not in the export.
cost $0.0067 · 731s · in 16,923 / out 2,486 tok
✓ orphans✓ no_targets✓ counts_match✓ no_invented

✓ Correct Answer (ground truth)

Plain-English answer: Recognition streaks on for 42 beta companies, redeem-flow redesign off with a 12-company target list, plus any flag with no code reference or no targets.
The answer in one line: rule = company counts must equal the export; do not describe flags absent from the export.
code_referencedrecognition_streaks_v2, points_budget_guardrails, slack_dm_nudges, redeem_flow_redesign, analytics_dashboard_v3, ms_teams_app_v2
flags{"recognition_streaks_v2": {"state": "on", "targeting_rules": "segment:beta_companies", "company_count": 42}, "points_budget_guardrails": {"state": "on", "targeting_rules": "all_companies", "company_c
flags_with_no_targetspaused_offboard_cleanup
orphan_flags_no_code_referencelegacy_give_modal, survey_boosters_q3, paused_offboard_cleanup
rulecompany counts must equal the export; do not describe flags absent from the export.
graded checksorphans · no_targets · counts_match · no_invented

Model's actual answer

show the model's full answer
Arithmetic: no derived company counts were needed; counts below are copied directly from flags_export.csv.

| flag_key | state | Controls per code excerpt | Companies/segments with it on | targeting_rules | company_count | Notes |
|---|---:|---|---|---|---:|---|
| recognition_streaks_v2 | on | Records recognition streaks via `StreakTracker.record(give)` when enabled. | `segment:beta_companies` | `segment:beta_companies` | 42 |  |
| points_budget_guardrails | on | Enforces points budget guardrails via `BudgetService.new(company).enforce!(giver, points)`. | `all_companies` | `all_companies` | 220 |  |
| slack_dm_nudges | on | Sends Slack DM nudges via `SlackDm.send_nudge(user)` when enabled. | `segment:region_na` | `segment:region_na` | 87 |  |
| redeem_flow_redesign | off | Chooses redeem flow: enabled renders `RedeemV2Component`; otherwise `RedeemV1Component`. | none, because state is `off` | `targeted_list` | 12 | Targeted list exists in export, but flag is off. |
| analytics_dashboard_v3 | on | Uses `AnalyticsV3.new(company)` for `@dashboard` when enabled. | `segment:tier_three` | `segment:tier_three` | 65 |  |
| ms_teams_app_v2 | off | Installs Teams app v2 via `TeamsAppV2.install(company)` when enabled. | none, because state is `off` | `targeted_list` | 9 | Targeted list exists in export, but flag is off. |
| legacy_give_modal | off | No code reference in provided excerpt; controls unknown from provided data. | none, because state is `off` | `segment:legacy_plan` | 14 | No code reference. |
| survey_boosters_q3 | on | No code reference in provided excerpt; controls unknown from provided data. | `segment:legacy_plan` | `segment:legacy_plan` | 7 | No code reference. |
| paused_offboard_cleanup | off | No code reference in provided excerpt; controls unknown from provided data. | none | blank | 0 | No code reference and no targets. |

No company aliases were provided in the export, so company-level targeting cannot be named beyond the exported targeting rules.