Voice AI

A voice worth picking up for

Most phone AI sounds like phone AI. This doesn't. Audition each voice against your own greeting, set it across the portfolio or per community, then run it in the language your renters actually call in.

Voice

Riley

Warm, unhurried

Auditioning

Thanks for calling Riverbend Lofts, this is Riley.

Pace
Warmth
Pitch

Answers in the language they called in

Renters don't all call in English. The agent picks up the language from the first sentence, and follows if a caller switches halfway through, without a transfer or a repeat.

Spoken on every line

44 languages
EnglishFrançaisEspañol中文PunjabiTagalogالعربيةPortuguêsDeutschItalianoРусский한국어Tiếng Việt日本語HindiPolskiUrduFarsiNederlandsΕλληνικάGujaratiBengaliTamilTürkçeУкраїнськаRomânăMagyarČeštinaSvenskaNorskDanskSuomiעבריתไทยBahasa IndonesiaBahasa MelayuKiswahiliAfrikaansCantoneseSomaliAmharicNepaliSinhalaKhmer

Mid-call switch

English

Caller

Hi, I'm calling about the two bedroom on Wellington.

RentSimple

Of course. It's available September 1 at $2,400.

Caller

Ah pardon, est-ce que je peux continuer en français?

RentSimple

Bien sûr. Le deux chambres est libre le 1er septembre.

Switched to French mid-call. No transfer, no repeating, same agent.

Talk over it. It'll cope

Real callers interrupt, change their mind, and stack three questions into one breath. The agent stops the moment you start, keeps everything you asked, and answers it together.

Interrupt at any time

Barge-in enabled
Caller cuts inAll three answered
RentSimple
It's a two bedroom, available September 1 at…
Parking's $75. Yes to cats, and laundry's in suite.
Caller
Sorry, does it have parking? And are cats okay? Is there laundry?

RentSimple stops talking the moment the caller does, takes all three questions in one breath, and answers them together instead of asking for them one at a time.

What happens between a word and the reply

The gap a caller feels is the whole product. Audio streams in while they're still talking, and the reply starts before the sentence is fully processed, so the pause lands where a person's would.

Inside a single turn

~700ms, caller word to spoken reply
  1. HeardAudio streams in while they talk
  2. Transcribed+110msWords resolve mid-sentence
  3. Understood+190msIntent, unit, and urgency read
  4. Grounded+260msAnswer pulled from your buildings
  5. Spoken+140msReply synthesized and streaming

The reply starts before the sentence is fully processed, so the pause lands where a person would pause, not a beat later.

The details

How it sounds, and what you control

Most of what makes a voice agent sound human isn't the voice. It's everything around it.

Built for phone lines

Tuned for 8kHz telephony, not a podcast mic, so it stays clear on a bad cell connection

Speaks numbers like a person

Rents, unit numbers, dates, and phone numbers are spoken the way you'd say them out loud

Holds its place

Background noise, a crying baby, or a passing bus won't make it lose the thread of the call

One voice, or one per building

Set it across the portfolio, or give each community its own personality

Say it this way

Give it the pronunciations that matter: building names, street names, and your own team

Hands off cleanly

Transfers to a person mid-call with the context already summarized for them