Voice

Talk to your agents like a phone call. Ask a question and get a spoken answer; ask for a post, an image, an email or a document and the finished card lands on your screen while you keep talking. Voice never publishes or sends — you tap the card.

Diese Seite ist in deiner Sprache noch nicht verfügbar — wir zeigen die englische Version.

Tap Voice and you're in a call. Ask a question, get a spoken answer. Ask for a post, an image, an email or a document — the agent reacts in its own words, and the finished card appears on your screen while you're still talking.

Two things happen at once on a Unyo call. The agent you're speaking to answers you immediately — it never goes quiet waiting on a tool. Behind it, the same agent's real toolset runs in the background and drops the result on your screen when it's ready. You keep the conversation; the work shows up.

Voice never publishes and never sends

This is deliberate, and it's enforced in code — not just asked of the model. Irreversible actions (publishing to Instagram, Facebook, LinkedIn or X, sending an email, posting to Slack, deleting a calendar event) cannot fire from a call. The agent prepares the draft; you tap Publish or Send on the card yourself. If you ask it to publish out loud, it will tell you — in one short sentence — to tap the button.

Start a call

  1. Open a conversation with any agent.
  2. Leave the message box empty — the Voice button appears where Send normally sits.
  3. Tap it. You'll see Connecting…, then Listening…
  4. Start talking. There's no greeting and no menu — you open the conversation, like calling a colleague who picked up.

You can start a call from a brand-new chat. The conversation is created for you and appears in your sidebar the moment you actually say something — and it's cleaned up if you hang up without speaking.

Which agents you can call

Nine agents take calls — Ashley, Maya, Alex, Riley, Tyron, Sam, Lucy, Blake and Ema. Each has its own way of speaking.

Steve, the website builder, is text-only. Building a site is a code-and-preview job, not a spoken one.

What you can ask for

Just talk

Anything that isn't a deliverable gets answered out loud, right away — a question, an opinion, advice, a brainstorm, an explanation, a story, a joke. The agent replies in your language, usually in a sentence or two, because it's being read aloud rather than displayed.

What should I focus on this week given what we talked about yesterday?

Give me three angles for a launch post about the new menu — just talk me through them.

The agent already knows you. It uses your first name, your company, your Neural Core entries, your contacts and your products — the exact same context your text agents have. You're not starting from zero because you switched to voice.

Ask for a deliverable and watch it land

Ask for something draftable and the card comes to you.

Make me an Instagram post about our Saturday crêpe special.

Generate an image of a sunset over the terrace.

Draft an email to Paul about moving the meeting to Thursday.

Here's the sequence:

  1. You ask.
  2. The agent reacts naturally — in its own words, different every time. It won't read the caption and hashtags back to you; that's the card's job.
  3. A moment later, a preparing placeholder appears with a quiet progress sound.
  4. The real card replaces it, with a notification sound. Social post drafts, images, email drafts, documents, captions — the agent's actual tools, running for real.

You never stopped talking.

It'll ask if it needs to

If the agent is missing something, it asks you out loud — "what's the subject?" — instead of guessing or silently producing a half-thing. Answer it like you would on any call.

Change what you just got

Edits are first-class. Say what's wrong and a fresh card comes back.

Change the door to green.

Make the background orange instead.

Say no and mean it

Changed your mind mid-generation? Say "no", "cancel", "forget it", "never mind" — or just "nope". The card isn't hidden; it's deleted. It won't appear during the call and it won't appear after a reload either.

This only listens during the moment a card is actually being prepared, so a stray "no" in the middle of a normal conversation can't trigger it.

During the call

ControlWhat it does
Mute / unmuteCuts your mic, with a distinct sound each way.
Red hang-up buttonEnds the call.
Escape keyAlso ends the call.
The orbReacts to both voices. Status reads Connecting… / Listening… / {agent} is thinking… / {agent} is speaking…
The + tools buttonGreyed out during a call — attachments and tools are a text-mode thing.

You can talk over the agent. Barge-in is filtered so its own voice coming out of your speakers doesn't cut it off mid-sentence, and echo cancellation runs on your side of the connection.

The transcript survives the refresh

Both sides of the conversation stream into the chat as live bubbles, word by word, while you speak. When the call ends, that transcript stays.

Reload the page and you see exactly what you heard: your spoken turns, the agent's spoken turns, and the card. Nothing more. No unspoken prose the agent never said out loud, no repeated partial sentences from the speech recognition refining itself, no agent avatar header cluttering a conversation you had by voice.

Languages

Set what Unyo listens for in Settings → Preferences → Voice language.

You get auto-detect plus 11 languages: English, French, Spanish, German, Italian, Portuguese, Turkish, Arabic, Chinese, Japanese and Russian. The setting is saved to your account — you set it once.

Auto-detect vs. picking a language

Picking your language is more precise than auto-detect, which understands several languages but commits to none. For English, French, Spanish, German, Italian, Portuguese, Japanese and Russian, choosing the language gives the recogniser a hard lock. Turkish, Arabic and Chinese are understood, but through the multilingual model rather than a single-language lock — so recognition is good, not laser-tuned. We'd rather tell you that than let a call fail.

The agent replies in whatever language you speak.

It knows your vocabulary

Generic speech recognition mangles company names. Yours doesn't have to be generic.

Every call rebuilds a recognition boost list from your live Neural Core — your company name, your product names, your contact names, plus Unyo and the agent names. So "send it to Amélie at Nordwerk" comes out as Amélie and Nordwerk, not as whatever the model guessed.

It's rebuilt fresh on each call, which means renaming a product or deleting a contact just works — nothing to re-sync. Only names are ever used for this. Never an email address, never a token, never anything sensitive.

Make this work harder for you

The more complete your Neural Core is — company, products, contacts — the better voice hears you. It's the same entries your text agents read, so filling it in pays off twice.

Products, by name

Mention a saved product by name in an image request and the agent generates using its reference photo and your saved style preferences.

Make me a photo of the Aurora tote on a café table in the morning light.

This only triggers on an actual name match in what you said — so a generic request never comes back with a surprise product in it.

What voice deliberately won't do

Enforced limits, not missing features:

  • No publishing. No sending. No posting to Slack. Prepare it by voice, commit it with a tap. Seventeen irreversible actions are gated this way.
  • No web research during a call. Search is switched off in the background executor — the agent answers you from what it knows rather than making you wait on a lookup mid-sentence.
  • No tools panel. The + button is disabled for the duration of the call.
  • No Steve. The website builder is text-only.

If a confirmable action ever does come up in a call, you get a Yes / No card above the orb that you have to click. Nothing commits itself.

FAQ

Can the agent publish a post or send an email while I'm on the call?

No — and this isn't a setting you can turn off. Irreversible actions are blocked from firing unattended, and the agent is instructed to tell you to tap the button on the card instead. It prepares; you commit. Ask it to publish and it'll say so out loud in one short sentence.

What happens to the conversation after I hang up?

It stays in your sidebar like any other chat. Reload the page and you'll see your spoken turns, the agent's spoken turns and any card it made — matching what you actually heard. If you start a call from a fresh chat and hang up without saying a word, the empty conversation is cleaned up for you.

I asked for a post and no card appeared. Why?

Usually because the agent didn't read your turn as a request to create something right now — for example if it offered rather than agreed, or asked you a clarifying question first. Try again with a direct instruction ("make me a post about X") rather than a question ("could you do a post?"). If it asked you something out loud, answer that first — the card follows.

Can I use voice with any agent?

Nine of the ten: Ashley, Maya, Alex, Riley, Tyron, Sam, Lucy, Blake and Ema. Steve, the website builder, is text-only.

Which language should I choose?

Pick your actual language rather than auto-detect if it's one of English, French, Spanish, German, Italian, Portuguese, Japanese or Russian — you get a hard recognition lock. Turkish, Arabic and Chinese work through the multilingual model. Auto-detect is the right pick only if you genuinely switch languages mid-call.

Does voice see the same information as the chat?

Yes. The call is built with the same context your text agents get: your name, your company, your Neural Core entries, your contacts and your products. Telling Maya something by voice isn't a separate memory — it's the same Neural Core Ashley reads next time.

Troubleshooting

"Connection problem" — Check your mic permission in the browser, then try again. If you're on a Steve conversation, that's expected: Steve doesn't take calls.

The agent can't hear me — Confirm you're not muted (the mic button plays a distinct sound each way), and that the browser has permission to use your microphone. If your OS is routing audio to a different input, the orb won't react when you speak — that's the tell.

It keeps mishearing a name — Add it to your Neural Core as a contact, product or your company name. The boost list is rebuilt on your next call, so hang up and call back after adding it.

It answered in the wrong language — Set Settings → Preferences → Voice language to your language instead of auto-detect.

The agent cut itself off / I keep interrupting it — Use headphones. Echo cancellation handles speakers well, but headphones remove the problem entirely.


Related: How agents work · The Neural Core · Social drafts · Email basics