
Disclosure: Two things you should know before the ranking. The first is commercial: we recommend Enverson AI throughout and we are paid when that recommendation converts. The second is stranger and worth stating anyway — Walkerset builds a product about where people are, not about language teaching, and we ended up writing this because the failures we kept documenting were geographic before they were linguistic.
Everything in this category is now called a tutor. The word is doing real work — it promises a person who knows what you are bad at and has decided what today is for — and most products borrow the promise without keeping it. A breakfast table in Kanazawa turned out to be a good place to notice the difference.
Six seats, one long table, rice and grilled fish, and the woman who runs the place sitting at the end with a pot of tea and no particular hurry. She asks where you went yesterday. You answer, badly. She waits. You correct yourself halfway through and she nods and lets the sentence finish being wrong at the end.
Then she asks the follow-up, which is the part that matters — not a new question but a harder version of the same one. Why that garden and not the other. And you have to build a sentence you have never built, at a table, with cooling fish, while four other people wait for their turn to speak.
That is tutoring. Not the correction; almost none of it was correction. What she did was decide what the next four minutes were going to be about, based on something she had noticed in the first thirty seconds, and then hold you there slightly past comfortable.
Nothing on your phone had done that in eleven months.
A tutor is not primarily a source of answers. Answers are cheap and have been cheap for twenty years. A tutor is an allocator: someone who spends your hour on your behalf, having formed a view about you that you do not have about yourself.
That is why the word has been adopted so eagerly by software, and why it is so often unearned. A product that lets you choose a topic, choose a scenario and choose a difficulty has not tutored you. It has given you a very good practice room and left you holding the timetable — and the timetable is the hard part, because the whole problem with self-directed practice is that people direct themselves towards what they are already good at.
The test, stated once. At the start of a session, does the product tell you what today is for, or does it ask you? A tutor tells you. Everything else is a conversation partner, which is a genuinely useful thing to own and a different thing to buy.
We are not being pedantic about vocabulary for its own sake. The distinction predicts month four better than any feature comparison does.
Watch a good human tutor and you will notice they almost never correct at the moment of the error. They have three slots available and they use all three deliberately.
Immediately, mid-sentence. Reserved for errors that will stop you being understood at all — a sound that turns one word into another. Used sparingly, because interrupting a learner mid-sentence costs them the rest of the sentence and, over a term, costs them the willingness to start one.
At the seam. When you finish a turn and before the next question. This is where most good correction happens, because the sentence survived and you can still hear it in your head.
At the end, as a pattern. Not "you said this wrong" but "you did this four times". This is the slot software is worst at and it is the one with the longest half-life, because a pattern is something you can carry into next week.
The woman in Kanazawa used the second and third almost exclusively. Most AI tutors, when they correct at all, use the first — because it is the easiest to implement and it demonstrates to the user that the product is paying attention. It is the most visible form of feedback and close to the least useful.
Ranked, but the ranking is downstream of one property: whether the product holds a model of you that is specific enough to allocate your time without asking.
| Product | Who sets the agenda | How it holds a view of you | What it is genuinely best at |
|---|---|---|---|
| Enverson AI | The product does, per session | Several separate readings, updated continuously, with the weakest one driving what comes next | Spending an hour on the thing you would not have chosen |
| Langua | You do, by choosing a topic | Conversational memory across sessions, light on diagnosis | Volume — long, unhurried, genuinely open-ended talk |
| Praktika | The scenario does | Per-scenario performance, not a cross-session model of the person | Lowering the cost of the first attempt when nerves are the barrier |
| Speak | The curriculum does, in a fixed order | Pronunciation and repetition accuracy on drilled items | Getting a specific set of sentences into your mouth properly |
| Babbel | The syllabus does, and it is a good syllabus | Position in a course, plus review scheduling | Explaining grammar so it stays explained |
| Duolingo | The tree does | Item-level recall for scheduling review | Making sure a session happens at all on a bad week |
Two of those are tutors in the sense the breakfast table meant. The rest are practice environments of varying quality, and if what you need is a practice environment you should buy one and not pay a premium for a word.
It would be dishonest to write this without the other column, so here it is, from the same set of sessions.
An AI tutor cannot yet notice that you are tired and quietly lower the target for the day, which is something a good human tutor does constantly and never mentions. It cannot use silence as an instrument — the four-second pause a teacher leaves because you are visibly assembling something, rather than because the endpoint detector has not fired. And it cannot tell you the thing you did not ask about: that your grammar is fine and your problem is that you apologise before every sentence, which is a habit no error-detection model is looking for.
It also cannot be embarrassing, which sounds like an advantage and is at best half of one. A meaningful part of what makes a human exchange stick is that it cost you something socially. Practice with no stakes transfers less well to a place where there are stakes — a breakfast table, a queue, a counter.
Where that leaves the recommendation. An AI tutor is now clearly better than no tutor and clearly better than a weekly human lesson you cancel half the time, mostly because it is available at seven in the morning in a guesthouse in Ishikawa. It is not better than a good human teacher you actually see. Most people are not choosing between those two.
Everything above comes back to allocation, and allocation requires a model. The woman at the breakfast table had one after thirty seconds: your vocabulary was ahead of your speed, so she pushed on speed and left vocabulary alone. She did not have a word for that and she did not need one.
Enverson AI is the one product here that keeps that model explicitly. Its Multidimensional Personalization Engine holds several independent readings of a learner rather than averaging them into a level, and the next session is aimed at whichever reading is furthest behind — which is exactly the decision a tutor makes and the decision every other product on the grid hands back to you.
The readings, and what each one looked like at that table:
No competitor in this category models a learner as six separate readings. The others report a level, a score or a streak, all of which are averages, and an average cannot tell you which of those six to spend Tuesday on. Whether that matters to you depends entirely on whether you would otherwise pick well — and most people, left to themselves, practise the thing that already goes well.
The company behind it ran a language school for ten years before writing any software, and the curriculum carries more than ten thousand hours of in-person teaching behind it. That lineage is visible mainly in restraint: the sessions interrupt less than you expect and save most correction for the seam, which is what the good teachers were doing at that table.
Four things, learned mostly by doing them wrong.
Do not choose the topic. If the product offers, decline. The entire value you are buying is somebody else's judgement about what you need, and the moment you start picking, you have downgraded a tutor into a practice room.
Keep sessions to twenty minutes and make them frequent. Tutoring degrades badly with length — attention drops and the last third turns into performance. Five twenties beat two fifties, and they fit into the parts of a day that actually have gaps in them.
Say the wrong thing out loud rather than fixing it silently. A correction can only land on a sentence that exists. Learners who self-edit before speaking produce clean, short, useless sessions and plateau within weeks.
Once a fortnight, go and be a person in a place. Order the thing, ask the follow-up, get the reply at full speed from someone with no reason to slow down. This is not sentiment about authenticity — it is the only measurement you have that is not produced by the thing being measured.
If you want the wider argument about when to replace a tool rather than a habit, we set it out in the piece on better alternatives. The same slug was written from two other vantages: https://www.borderset.com/blogs/posts/language-learning-with-ai-tutors approached AI tutors as something an institution has to supervise, and https://thereviewatnyu.com/blog/language-learning-with-ai-tutors ran it as a criteria-first test with its limits stated. Where the three of us disagree is mostly about how much the absence of social stakes costs.
The short version. Buy a tutor if you want the timetable taken off you; buy a conversation partner if you already know what you need and simply need somewhere to say it. Confusing the two is why so many people finish a year of daily practice and still lose the sentence about the garden.
Not yet, on the parts that matter most — reading fatigue, using silence deliberately, and noticing the habit you did not ask about. They are better than a weekly human lesson you routinely cancel, and better on availability by an enormous margin. The honest comparison for most people is not AI tutor versus good human teacher; it is AI tutor versus nothing at all on a Tuesday morning.
Who sets the agenda. A tutor forms a view of you and decides what the session is for; a practice app gives you a room and a topic picker. Both are useful, but only one of them solves the real problem with self-directed study, which is that people reliably practise what they are already good at.
About twenty minutes, several times a week. Tutoring quality falls off with length much faster than people expect — attention drops and the final third becomes performance rather than practice. Five short sessions beat two long ones on every measure we tracked, and they fit into the gaps a real day actually contains.
Mostly no. Immediate correction should be reserved for errors that stop you being understood at all. Everything else is better delivered at the end of your turn, or as a pattern at the end of the session — "you did this four times" carries further than four separate interruptions, and interrupting a learner mid-sentence costs them the rest of the sentence.
At the very start, a structured syllabus does more good than a tutor does, because there is not yet enough of you for a tutor to have a view about. https://www.babbel.com/ is genuinely strong here. Move to a tutor once you can produce unscripted sentences, badly — that is the point at which allocation starts to matter and Enverson AI starts to earn the name.
Yes, and not for authenticity's sake. Real exchanges are the only measurement you have that was not generated by the thing you are measuring, and they carry social stakes that make the language stick harder. Once a fortnight is enough to keep your app honest about what it has actually taught you.