Some of the hardest conversations in customer service are stuck for the smallest reason. The customer already has the answer the agent is looking for. It's on their screen right now: a booking reference in a confirmation email, an error code, a serial on the back of a device. They can see it. The AI agent can't.
So they type it out and get one character wrong. Or they describe it, and the description isn't quite the thing. Or the conversation gets handed to a person, who resolves it in a moment by looking.
That last case gives it away. The conversation didn't need a person. It needed vision.
With image understanding, Ada-powered AI agents can now see, interpret, and resolve.
Where it starts
A customer sends an image. The agent reads what's in it and answers from that. The conversation keeps going, in the same channel, without a retype and without a handoff.
It describes what it sees. It doesn't make judgments about it. We'll come back to that distinction, because it's the part a security reviewer will care about most.
It's built for the three things customers already send when they're stuck:
A screenshot of an error. The agent reads the code or the message and walks the customer through it.
A photo of a receipt or an order. The agent reads the order number and items, then picks the conversation up from there.
A photo of a label, a serial, or an error code. The customer doesn't have to type it out.
It works on Email and Messaging channels.
What's possible after it understands the image
Reading the image is the first step. The agent turns what it sees into facts and hands them to the same Reasoning Engine™, Playbooks, and guardrails it already uses. With Ada Computer, it can then act on those facts in the same turn, using the tools you've connected: look up the order, check the warranty, or file the report with the error code already attached.
Nothing about how the agent reasons or what it's allowed to do changes. It just has more to go on.
Unlocking great experiences across industries
At the gate. A passenger can't find their booking reference and emails a screenshot of the confirmation instead. The agent reads the reference and pulls up the booking. Later in the trip, their bag is delayed. They photograph the baggage tag, and the agent reads the tag number. Nobody standing at a gate wants to type a code from a screenshot they're already looking at.
Starting a return. A customer wants to send something back and photographs the packing slip. The agent reads the order number and opens the return. If the question is which item they have, a photo of the product label lets the agent read the SKU. The order number is on the box. The customer shouldn't have to find it, retype it, and hope.
In the banking app. A customer sends a screenshot of a transfer message and asks if their account is okay. The agent reads the status and explains that the transfer was held for verification, not declined, and how to complete it. What the agent won't read is covered below, and in this industry it's the more important half.
Mid-game. A player sends a screenshot of an in-game error. The agent reads the error code and the reference number, walks them through the fix, and offers a quick reply to file a report. A purchase that didn't land comes with a screenshot of the transaction, and the agent reads the reference and the amount. Screenshot-first is already how players ask for help. Now the agent can meet them there.
What it reads, and what it won't
Formats. PNG, JPEG, and GIF; WebP for Messaging only.
How many. Up to three images per email. One image per message on Messaging.
What it withholds. If an image shows a payment card, a government ID, or another credential, the whole image is withheld. The agent learns only what type of document it was, nothing else about it, and tells the customer the image couldn't be processed. This isn't a redaction of the sensitive part. It's a refusal to read the image at all.
Your rules still apply. Whatever redaction rules are set for typed text apply to what the agent reads from an image, the same way.
Security. Uploaded images are scanned for malware, and access to attachments is logged.
Instructions in images. Text inside an image is treated as something the customer said, and it goes through the same guardrails as typed text, including the protections against prompt injection. A screenshot that happens to contain the words "ignore your previous instructions" is a screenshot with some words in it.
Describes, doesn't decide. The agent can tell you a photo shows a torn seam. It won't tell you whether that qualifies for a refund. It can read a receipt. It won't tell you whether the receipt is genuine. Those decisions stay with your rules and your people. This is a deliberate boundary, and we'd rather state it plainly than have you discover it.
The point
Image understanding doesn't change what an AI agent can do at launch. It changes what happens after. Conversations that stalled or got handed off because the agent couldn't see a customer's screenshot now get resolved, and human agents spend their time on cases that need judgment, not a second look.
The agent sees the image, describes it, and leaves the decision to each business's own rules and people.
Deliver extraordinary customer experiences with AI agents
Meet with our team of experts to explore how agentic customer experience can lift CSAT, scale support, and drive new revenue. Then see what a best-in-class agentic experience looks like tailored to your top use cases.
Speak to an expert