Support is full of photographs. A cracked mug, a shipping label, a screenshot of an error, the wrong colour hoodie. Most AI support tools cannot see any of it: the customer attaches a photo, the bot answers the caption, and a human has to open the thread anyway to look.
What it does
Arbyn looks at the photo. Every inbound image is read before the agent classifies what the conversation is even about, so what it saw shapes the reply and the action, not just the wording.
It also survives a bare photo with no words at all, which is how people actually complain. A shopper who sends one picture of a smashed box and nothing else is not an empty message, and Arbyn does not treat it as one.
How store owners use it
- Damaged on arrival. The customer photographs the box. Arbyn reads the damage, matches it to the order, and puts the refund in front of you with the picture attached.
- Wrong item shipped. One photo tells Arbyn what actually arrived. It compares that to what was ordered instead of asking the customer to describe it.
- A receipt, a label, a screenshot. Arbyn reads the text in the image, so a customer who photographs a label rather than typing it out is not made to type it out.
When a message carries an image, Arbyn looks at it before it decides anything. The intent, the draft and the action are all made knowing what was in the picture.
What it costs
Nothing extra. Reading a photo is part of answering the conversation, and conversations are what the plan covers.
It is included on all three plans, Starter, Growth and Agent. There is no per-image fee and no vision add-on, because a support tool that charges you more when a customer attaches a picture is charging you for the thing you most wanted it to do.
How to enable it
There is nothing to switch on. If a message has an image, Arbyn reads it.
- Install Arbyn. Vision runs automatically on any inbound message that carries an image.
- A shopper attaches a photo, on email or on the chat widget.
- Arbyn reads it before it decides what the conversation is about, so the picture shapes the answer.
- What it saw is written into the conversation record, so you can check what Arbyn was looking at when it decided.
The fine print
Arbyn describes an image, it does not identify people. If the vision step fails, it fails open: the conversation continues on the text alone rather than stalling. What Arbyn saw is recorded on the conversation, so an escalation review can always answer the question that matters, which is what was it looking at when it decided that.