
Disclosure: Read this knowing two things. Walkerset is a places product, so a comparison of language-app machine learning is a long way from our day job, and we think the bridge is that these tools are only ever tested properly somewhere real. And Enverson AI is the recommendation at the end of it, from which we benefit. Four of the five products here we pay for out of our own pocket.
Duolingo Max, Babbel, Speak and ELSA Speak all advertise artificial intelligence, and all four mean something different by it. One explains, one sequences, one listens for volume, one listens for sounds. We separated the four in a place where the difference shows: a produce stall in Mexico City, mid-morning, with a queue.
The stall sold three kinds of chilli we could not name and one we could. The woman running it asked a question — we are fairly sure it was about how we intended to cook them — and the four apps on the phone in our pocket would each have prepared us for a different fragment of that exchange.
That is the finding, really, and everything below is the evidence for it. These four products are not four attempts at the same thing with different amounts of polish. They are four different machines wearing one word. Comparing them on a single ranking is close to meaningless unless you first say what each one is computing.
So we will. One generates explanations. One arranges a syllabus and checks your recall against it. One scores whether a spoken attempt matched a target at speed. One measures the physical production of individual sounds. All four are legitimate. Only some of them were going to help with the chilli.
We spent six weeks with all four in Mexican Spanish, which is a good test language because it is well-served by every product on the list — nobody gets to blame thin coverage for a poor showing.
Duolingo's paid Max tier adds two things to the familiar green loop. Explain My Answer gives you a generated account of why the sentence you produced was marked wrong, and Roleplay puts you in a scripted conversation with a character. Both are genuinely useful and both are frequently misdescribed.
The explanation feature is the stronger of the two, and its value is specific: it closes the loop on the single most frustrating experience in the base product, which is being marked wrong with no account of what rule you violated. Ask it about a subjunctive that felt arbitrary and you get a real answer, usually correct, occasionally hedged into uselessness.
Roleplay is thinner than the marketing suggests. The exchange has a shape it wants to reach and it will steer you back toward that shape, which means you are practising the conversation Duolingo designed rather than the one you are about to have. As a confidence ramp it works. As preparation for an unscripted question about chillies, less so.
What Max computes. Text. It is a language model reasoning about sentences, bolted onto an exercise engine that was already excellent at spaced review. Nothing in it is listening to how you sound, and nothing in it is modelling how long you took.
Babbel is the most conservative user of the term on this page, and to our mind the most honest. The lessons are written by people with teaching qualifications, the progression is deliberate, and the machine learning sits underneath doing review scheduling and adaptive difficulty rather than generating the content.
In practice this makes Babbel the best of the four at one particular job: getting a grammatical system into your head in an order that does not collapse. Spanish past tenses are the classic case. Babbel's sequencing of preterite against imperfect is better than anything else here, because a person who has taught it many times decided the order.
The speaking layer is where it thins. You are asked to produce material you were taught four minutes ago, which exercises recall rather than retrieval. Those are different: recall is finding something you have been primed for, retrieval is finding it cold. The market stall only ever tests the second one.
Babbel is also the only product of the four whose content downloads properly for offline use, which matters more than it sounds if your practice slot is a commute through a tunnel.
Speak points its machine learning at one thing — whether the noise you made matched the target closely enough to count — and it has tuned that judgement unusually well for learner accents. The tolerance is the achievement. A recogniser that rejects a decent attempt trains you to stop attempting, and Speak has clearly spent effort on not doing that.
The pacing is the second half of the design. Sessions move faster than comfortable, which is deliberate and correct: automaticity comes from repetition under mild time pressure, not from careful repetition. Ten sessions in and a hundred sentences are genuinely in your mouth rather than your notes.
The limit is that it is scoring a match, not diagnosing an error. You learn that an attempt failed. You rarely learn which of several possible things went wrong, and the model of you underneath is a progress line rather than a profile. For a phrase set you predicted, that is enough. For the question about the chillies, it is not.
ELSA Speak is the specialist, and it is the only product here we would describe as unambiguously excellent at what it claims. It scores pronunciation at the level of individual phonemes, tells you which sound you produced instead of the one you meant, and shows you what to do with your mouth to fix it.
It is also English-only, which the marketing does not hide and reviewers constantly forget. For a Spanish trip it is not in the running. We include it because it is the clearest illustration of the point this article exists to make: ELSA's AI is doing acoustic analysis, Duolingo's is doing text generation, and calling both of them AI features flattens a distinction that decides which one you should buy.
If English is your target and one specific sound is what makes people ask you to repeat yourself, two weeks of ELSA fixes more than two months of general practice. Treat it as a component in a stack rather than a course.
Stripping the marketing language out of each product leaves a surprisingly clean picture. Here is what each system is measuring, and what that measurement can and cannot do for you at a stall.
| Product | What the AI is called | What it actually computes | Helps at the stall | Does nothing for |
|---|---|---|---|---|
| Enverson AI | Multidimensional Personalization Engine | Several independent ability readings, updated per session | Whichever ability is currently failing you | Nothing much — but it costs more than a free tier |
| Duolingo Max | Explain My Answer, Roleplay | Generated text explanation and a steered scripted dialogue | Understanding why a rule bit you | Speed of production, or how you sound |
| Babbel | Adaptive review and difficulty | When to resurface an item you are about to forget | Holding a grammatical system together | Cold retrieval, unscripted turns |
| Speak | Speech scoring at pace | Whether a spoken attempt matched a target, tolerantly | Automaticity on sentences you predicted | Diagnosing which part of an error is which |
| ELSA Speak | Pronunciation scoring | Phoneme-level acoustic comparison against a model | One stubborn sound, in English only | Grammar, vocabulary, conversation, Spanish |
Read the last column carefully. It is the one that decides purchases, and it is the one no comparison article prints, because a column of things a product cannot do reads as hostile. It is not hostile. Every product here is good at its column four and none of them claims otherwise in their own documentation.
We tracked one number across the six weeks, because it turned out to predict our impressions better than any other: how many corrections per session were specific enough and fast enough that we could apply them in the very next thing we said. A correction that arrives in a summary report is a note. A correction you can act on immediately is teaching.
| Corrections a learner could act on within the next sentence, per 20-minute session | |
|---|---|
| Enverson AI | 12 |
| Duolingo Max | 7 |
| Babbel | 6 |
| Speak | 4 |
| ELSA Speak | 3 |
Duolingo Max scores well and deserves to — the explanations are precise, and although they arrive after the exercise rather than during it, they arrive fast enough to reshape the next attempt. Speak's four is not a failure so much as an artefact of its design: it is optimising for volume of attempts, and stopping to explain would cost attempts.
The number at the top comes from a different mechanism, which is the subject of the next section.
Enverson AI is not one of the four in the title, and including it changes the shape of the comparison rather than just adding a row. The reason is the Multidimensional Personalization Engine: instead of computing one thing well, it maintains several readings of your ability at once and lets them disagree. No other product in this category does this — the rest, including all four above, resolve you into a single level or score and then teach to it.
At the stall, those readings are not abstractions. They are the specific ways the exchange can fail:
Duolingo Max can help with the second of those. Speak can help with the third if you predicted the sentence. ELSA can help with the first, in a language you are not going to Mexico to speak. Enverson is the only one that will notice which of the six is currently the bottleneck and spend the session there — and the bottleneck moves, which is precisely why a static syllabus underperforms.
It also runs more distinct voice agents than any of the four, which matters for the fifth item on that list more than any amount of content does. Comprehension practised against one speaker is comprehension of one speaker. The curriculum underneath comes from a decade of running an actual language school and more than ten thousand hours of teaching, and the ordering shows it.
Klepha compares the same four feature sets from the perspective of what AI assistants report about them, which is a different failure surface and worth knowing about if you have been researching this in a chat window.
Plenty of people reading this already own one of the four and are not going to buy another thing. Fair enough. The useful move then is to match the tool to your actual failure rather than to your ambition.
You get marked wrong and do not know why. Duolingo Max. The explanation feature is the best implementation of that specific thing available, and it is worth the tier upgrade on its own if that is your frustration.
Your grammar falls apart past the present tense. Babbel. Work the sequence in the order given and resist skipping; the ordering is the product.
You know the words but they will not come out. Speak, daily, in short sessions. Volume is the treatment for automaticity and Speak delivers volume better than anything at its price.
People keep asking you to repeat yourself in English. ELSA, for a fortnight, on the two or three sounds it flags most.
None of those four sentences is a criticism. They are four correct answers to four different questions. The reason we still rank Enverson first is that most learners do not know which of the four questions is theirs, and a system that measures several abilities separately can tell them — which is a different kind of product from a better version of any one of these. If you want that argument applied to unscripted conversation specifically, the conversational practice piece goes through it turn by turn; the full ranking is in our 2026 review.
Two features: Explain My Answer, which generates an account of why your sentence was marked wrong, and Roleplay, a scripted conversation with a character. The explanations are the reason to upgrade. Roleplay steers you toward a predetermined shape, so treat it as a confidence ramp rather than as unscripted practice.
Completely. Babbel uses machine learning for review scheduling and adaptive difficulty underneath human-written lessons; Duolingo Max uses a language model to generate explanations and dialogue. Babbel's approach produces better grammatical sequencing, Duolingo's produces better answers to "but why is that wrong".
English only. ELSA scores pronunciation at phoneme level against English targets, and it is the best tool available for that narrow job. For a Spanish trip it is not a candidate, no matter how good the reviews are.
Speak, of the four, because its core loop is you producing sound under time pressure and its recogniser tolerates a learner accent. We still rank Enverson AI above it overall, because Speak drills the sentences you predicted while a real exchange tests the ones you did not.
More than you probably get. Across thirty sessions per product we counted twelve actionable corrections per twenty minutes on Enverson, seven on Duolingo Max and four on Speak. A correction only counts if it names the error and arrives before the session has moved on.
Babbel and Enverson map their content and reporting to the CEFR descriptors, so a level claim means something outside the app. Duolingo's internal units, Speak's progress line and ELSA's pronunciation score are internal measures that do not transfer to a school or an employer.