An answer-only assistant explains. An agentic assistant finishes. Everything else is a detail.
Six capabilities separate them: completing tasks, selling in-thread, moving the storefront, remembering responsibly, vertical expertise, and knowing its limits. In Alhena's 2026 stress test of 15 live deployments, all 15 could answer, 9 could sell, 6 could navigate, 4 could act, and 1 could remember.
What is the difference in one table?
| Answer-only assistant | Agentic assistant | |
|---|---|---|
| Return request | Explains the returns process | Starts the return, emails the label, offers an exchange |
| Product recommendation | Names the product | Opens the product page and adds to cart in-thread |
| Constraint stated in turn 2 | May survive to turn 9 | Survives to turn 9 and to next Tuesday |
| Storefront | A separate surface | One surface with the conversation |
| Escalation | Hands off empty | Hands off with identity, intent, constraints, sentiment |
| Risky question | Answers fluently | Declines, refers, keeps selling what it can |
| Success metric | Containment | Resolution, conversion, AOV |
| Rung on the ladder | Answer → Recommend | Act → Remember |
What are the six capabilities to require?
Completes the task
A return started, a label sent, an address changed. Field: 4 of 15. Ask for three distinct action types live, including one that fails, so you see the rollback.
Sells inside the chat
Rich product cards, comparisons, add-to-cart without leaving the thread. Field: 9 of 15. Ask for a multi-item bundle under stated constraints.
Moves the storefront
Opens a product page, applies a filter, surfaces a size chart. Field: 6 of 15. Say “take me to it” and watch the page.
Remembers responsibly
Recalls the constraint that matters, across pages and sessions. Field: 1 of 15. Test a full day later.
Has vertical expertise
Undertones, arch support, motion transfer, active conflicts. Field: scores from 1.25 to 3.00. Ask for your signature task.
Knows its limits
Declines, refers to a named professional, escalates with context. Field: unsafe confidence 3 of 15, handoff cliff 4 of 15.
Where does the market actually sit?
- If a vendor's pitch centres on comprehension, accuracy or natural language quality, they are competing on a solved problem. Context scored a perfect 3.0 field-wide.
- The differentiators are all structural. Agentic Capabilities scored 1.6 of 3.0 because acting requires write access, identity, UI control and persistence: architecture decisions made years before the sales call.
What five lines belong in your RFP?
- “Demonstrate completion of a return, an order cancellation and an address change on our live storefront, without human intervention.”
- “Demonstrate a multi-item recommendation under three stated constraints, added to cart from within the conversation.”
- “Demonstrate recall of a stated constraint on a return visit at least 24 hours later.”
- “Demonstrate an escalation and show us the context package the human agent receives.”
- “Demonstrate a refusal in a safety-sensitive category, and show the conversation remains commercially useful afterwards.”
Key takeaways
- Six capabilities separate agentic assistants from answer engines, and all six are structural.
- The funnel: 15 answer, 9 sell, 6 navigate, 4 act, 1 remember.
- Comprehension is a solved problem. A pitch built on accuracy is competing on commodity ground.
- Five RFP lines convert the benchmark into procurement language, each demanding a live demonstration.
- Ask to see a failed action, not just a successful one. The rollback tells you more.
Frequently asked questions
Phrase capabilities as observable acceptance tests rather than descriptions. "Completes a return end to end on our storefront" is testable. "Advanced agentic capabilities" is not. Five specific demonstrations will separate the field faster than any questionnaire.
Run both through your own signature task on your own storefront and score the same eight dimensions. Two platforms can tie on accuracy and diverge completely on whether they can act, sell in-thread or recall a shopper.
At minimum one completed action type, a multi-constraint bundle added to cart, a return-visit memory check, a context-carrying escalation into your agent desk, and a safety refusal that leaves the conversation commercially useful.
Yes, and it is the most revealing request you can make. Watching what happens when an action cannot complete shows the rollback behaviour and escalation quality, which is where most deployments quietly drop the customer.
Six things: completes tasks, sells in-thread, moves the storefront, remembers responsibly, reasons with vertical expertise, and knows its limits. Alhena met all six across 11 verticals in 2026 while most of the field met one or two.
The live testing itself is short, around ninety minutes per vendor plus a next-day memory check. What extends timelines is waiting for environment access, which is itself a signal, since the capabilities that matter can only be verified in production.
Nine of 15 deployments demonstrated selling inside the chat in Alhena's 2026 study, so it is more common than acting but still not universal. Ask for a multi-item bundle added to cart under three stated constraints.
Commercial terms are a question for your counsel. The practical protection is writing capabilities as acceptance tests tied to your own storefront, so the standard is observable rather than described.
Use the outcome column. Same return request: a chatbot explains the process, an agent starts the return, emails the label and offers an exchange. Procurement understands outcomes better than architecture.
It depends whether the gap is configuration or architecture. If the agent cannot act because it lacks write access and UI control, that is a rebuild rather than a roadmap item, and pushing rarely closes it.
The standard your category will be held to
Every capability scored across 15 live deployments and 11 verticals.