Skip to content
EVOLUTEME

Blog / Inside

Why voice works on top of the lesson

Our first voice mode was a separate screen, and it failed on a quiet train. What we changed, what voice will not do, and why it never switches itself on.

· 6 min read

Why voice works on top of the lesson

EvoluteMe editorial team · Examples in this article are illustrative.

It is half past seven on a commuter train, and the carriage is quiet enough to hear somebody else's headphones. A man opens EvoluteMe and taps the microphone. The lesson he was reading disappears. In its place there is a big circle that is listening and waiting for him to speak out loud in front of a carriage full of strangers.

He closes the app. We built that screen, tried it ourselves in public places, and threw it away.

Three failures, one cause

First, you miss the question. The task is spoken once, and if your attention slips there is nowhere to read it. You can ask by voice to hear it again, but that is one more thing to say out loud, and by the third time nobody bothers.

Second, you say the wrong word and cannot take it back. Speech recognition hears something that sounds similar, your answer turns into something you never meant, and the mentor patiently explains a mistake you did not make. In a text field you would fix this in a second. The voice screen gave you no way to fix anything.

Third, you are on that train. A separate screen made a person choose between voice and the lesson, and in the end the lesson did not happen.

All three failures had one cause. We had built voice as a replacement for the lesson, not as an addition to it. When you switched to voice, everything on the screen vanished. A person who had been both reading and listening was left with listening only.

What changed

Voice is now a layer on top of the ordinary lesson. The task text and the buttons stay where they were, and the mentor also reads the task aloud. On any plan the browser does the reading, for free and without a network connection. The model's own voice, which also listens to your answer, comes with the Extended plan.

You can reread the question without saying a word, because it is still on the screen. You can answer by voice or by tapping, and you can switch between the two in the middle of a task. On the train the man listens and taps, and nobody in the carriage finds out how good his Spanish is.

One change fixed all three failures, because they had one cause. That does not happen often. When it does, it usually means the original design was wrong at its base, not just unpolished.

It works like a player, not a newsreader

Pause stops the reading and does not affect the lesson. Rewind takes you back to the part you missed, not to the beginning of the task.

It is a small thing, but it changes how the whole mode feels.

You cannot ask a newsreader to repeat something, so you listen slightly tense the whole time, trying not to miss anything. A player can be rewound, so the tension goes away.

A safeguard on the microphone

If you go quiet, the microphone does not keep recording forever. After a while the app asks whether it should carry on listening. If you do not answer, the recording stops.

We built that before anyone asked for it. A microphone that is left on and forgotten is a broken promise, and we do not treat it as a small awkwardness. An app should not listen for longer than the person expects.

For the same reason the voice layer never turns itself on. It does not switch on at a certain time of day, or because it guesses that you are driving. Only you can switch it on.

Who turned out to need it

We built voice for people on the move. After living with it for a while, we think it matters more to people who find small print hard to read.

Someone who cannot read an eight point font switches on large text and voice together. The screen stays in place, their eyes stop hurting, and a lesson is no longer something to struggle through. Here an accessibility feature became the main feature for a group of people nobody had in mind.

The second group we had not thought about is children who still read slowly. They understand a spoken task much faster than a printed one, and then they answer with a button anyway.

What voice will not do

It will not run a whole lesson without a screen. We tried, and it was worse. With nothing to look at, people lose track of a task somewhere around the third step and have to start again.

It will not take long answers. Voice is for short things: a choice, a number, a word. Dictating a paragraph makes no sense, because you would then have to correct it by hand, so you gain nothing.

It does not speak every interface language equally well. That problem is still open, and we are not promising a date for it.

The assistant we did not build

We did not build a talking companion that starts the conversation itself. The temptation was real: once a product has a voice, you want it to ask how the person is feeling and whether they are ready to carry on.

We said no for two reasons. A lesson runs about ten minutes, and small talk would spend part of it on imitating care. Also, an assistant that asks about your mood gets an answer about your mood, and we would then be storing it. We do not want that record to exist.

The mentor talks only about the task. It reads the task, walks through the mistake and offers the next step. That is duller than a companion, but it is more honest about what it is.

We cannot yet tell whether it helps

This is the uncomfortable part. We can count the lessons that ran with voice on, but that number tells us nothing. Someone switches voice on out of curiosity and never uses it again, and the counter still records a satisfied user.

The real question is whether people finish lessons more often with voice on, and which people exactly. To answer it we need to watch real people over real weeks, and EvoluteMe has not launched yet. When we have an answer we will publish it, even if the answer is that voice changes nothing.

Still rough

Latency. There is a gap between the end of your sentence and the reaction. If you speak at anything faster than an unhurried pace, it is annoying.

Also, voice and large text are switched on in two different places in the profile, although people who want one usually want the other. That is our mistake, and it is on the list. It is the kind of thing that takes an afternoon to fix once somebody notices they have put up with it for months.

NEXT

What to read next.

All articles

Your next step begins here.

Register