The $100,000 Tutor Test
A new study will ask how far one-to-one tutoring can take a child. The bigger question is which parts families actually need to buy.
Reed Hastings is funding an education study that will spend up to $100,000 per child to answer a simple question: “How much can a young brain learn under ideal conditions?”
Benjamin Somers, Modern Bloom’s co-founder and head of program, is leading its 2026 2-Sigma Study. The project advertises an August 2026 start for virtual one-to-one tutoring in math and English as children move from third to fourth grade. Each child gets a dedicated tutor for the full year.
Somers says the pilot will randomly sample students across two schools without selecting them by ability. It will compare their results with both a direct comparison group receiving normal classroom instruction and a second benchmark built from other schools in the same network or region.
I think this is a great idea. Most education studies have to work within normal school budgets. This one deliberately removes that constraint. Somers and Hastings want to find out what is possible first, then see if technology can reproduce it cheaply. Their hope is to turn a $100,000 education into something available to every child for “under $100.”
But a big result will not tell us which parts of that $100,000 package did the work. The tutor will choose work at the child’s level, keep the child working, and teach new material or correct misunderstandings. One-to-one tutoring provides all three. They do not all require one adult sitting beside one child.
Tutoring works. Bloom’s two-sigma number did not hold up.
In 1984, Benjamin Bloom reported that students taught by tutors performed two standard deviations better than students taught in a conventional classroom. In his words, “the average tutored student was above 98% of the students in the control class.”
The underlying experiments really did produce that result. But they were unusually favorable: 11 lessons over three weeks, entirely new topics, and immediate tests written around what had just been taught. The tutored students also received mastery checks, corrective feedback, extra time, and specially trained tutors. The studies showed what one intensive package could do on a narrow short-term test. They did not show that tutoring normally moves a child from average to the 98th percentile in school. (For more on Bloom’s limitations, see Paul von Hippel or Daisy Christodoulou.)
Broader tutoring research has never reproduced that number. In a meta-analysis of 96 randomized experiments, none reached two standard deviations; the average effect was 0.37. Tutoring works, but the evidence does not justify treating two sigma as its benchmark.
Modern Bloom is not treating two sigma as a forecast. The team knows that result is not typical. It is asking how large a gain can be achieved over a full year, on a broad test, when tutoring resources are not constrained.
That is what makes the study interesting. It is a stronger test of what the full tutoring package can do under ideal conditions. But it will still test the package as a whole. Whatever result it gets, the next question will be: what, exactly, did the tutor add?
A tutor bundles three advantages. Other systems can provide them too.
Parents are usually asked to compare products: human or AI tutor, school or homeschool, app or textbook. I think the better question is what changes for the child.
A great tutor does three jobs:
Chooses the right next work. The tutor repairs missing prerequisites and skips material the child already knows.
Keeps the child working. The tutor notices when the child starts idling, guessing, or waiting for the answer, and helps them stay focused.
Provides good instruction. The tutor explains new material, corrects misconceptions, and gives useful hints and feedback.
These jobs can come apart. Grouping children by current level can fix the first without adding more adults. An app may choose the right problem but do little when a child refuses to use it. A teacher may explain brilliantly but have no practical way to aim every lesson at 30 starting points.
More time won’t fix work that is too easy or too hard.
Children in the same grade do not all start in the same place. A fixed grade-level curriculum therefore misses in two directions: some children lack the prerequisites, while others learned the material months or years ago.
When the work is too hard, more of it does not help
In the Mindspark experiment in Delhi, students in the same grade spanned five or six grade levels. Some were offered Mindspark, a program that continually adjusted the material to what each child knew; others continued with normal instruction. The researchers used tests that included material well below grade level, specifically so they could measure progress among weaker students. After four and a half months, the estimated progress of the bottom third receiving normal instruction was near zero and statistically indistinguishable from no progress. Mindspark, on the other hand, produced gains across all starting levels. Its relative benefit was greatest for the weakest students because normal instruction was producing almost nothing for them.
A second Mindspark experiment in Rajasthan compared two versions of the same adaptive math software. One could assign exercises only from the child’s enrolled grade; the other could move back to earlier-grade material, such as basic number sense before fractions. Among children who started in the bottom quarter, access to earlier-grade material produced an estimated gain of 0.22 standard deviations.
The same idea worked without software. In a randomized experiment in 121 Kenyan schools, every school split its first-grade class in two. Some grouped children by starting achievement; others split them randomly. Grouping by level raised scores by about 0.14 standard deviations after 18 months, with gains in both the higher and lower groups. For lower-achieving children, the authors concluded that grouping helped teachers work at their level.
Together, these studies show that matching instruction to a child’s level can help - and that software or grouping can do it without a one-to-one tutor.
When the work is too easy, a child can spend months waiting
In a large study of high-achieving students, teachers tested what each child knew before a unit, then removed material they had already mastered. They removed as much as 50 percent of the regular curriculum for those students, and researchers detected no decline on out-of-level tests in reading, math computation, social studies, or spelling. This was an old study in volunteer districts, not proof that half of school is useless. But it shows how much material some advanced children may be required to repeat without measurable academic benefit.
Acceleration research points the same way. A 1992 review found that students allowed to move ahead finished 0.87 standard deviations above similarly able students who stayed with their age group. They performed about as well as the older students they joined. The evidence is mostly old and students were not usually assigned at random, so I would not treat 0.87 as a guaranteed causal effect. But it is hard to argue that making those children wait helped them.
Most classrooms include children who are ahead, behind, and in between
To see how wide the range can be inside one classroom, consider TIMSS, an international math assessment. It sorts performance into four bands: low, intermediate, high, and advanced. For fourth-graders, the low benchmark covers basic arithmetic and simple problems; the advanced benchmark requires multistep reasoning with fractions, geometry, and data.
A nationally representative analysis of US classrooms found that 23 percent of students in the typical fourth-grade math class scored at the low benchmark or below, while 14 percent reached advanced. Nearly seven in ten classrooms included children in every one of the four bands.
These performance bands do not identify the exact next lesson or prove that each child’s daily work was mismatched. But they show the basic problem with “grade level”: it is an age bracket, not a common starting point.
Tutoring does more than match the level. We do not know what each extra part is worth.
Matching the work to the child covers only the first job. A tutor also supplies attention, explanation, accountability, and perhaps a relationship. Few studies change one part at a time, but several offer clues.
One-to-one buys more attention. In one virtual reading trial, the authors’ preferred sample produced estimates of 0.12 standard deviations for one-to-one tutoring and 0.04 for two-to-one. The paper did not report a formal test comparing the formats, and the estimates are too imprecise to establish a difference. But an analysis of 16,629 recorded sessions found that one-to-one students received more personalized support, while two-to-one sessions left each child with more downtime.
Effective tutors do not have to be licensed teachers. Chicago’s Saga program used screened paraprofessionals, a structured curriculum, and frequent assessments. Two students worked with each tutor for 50 minutes a day. The pooled treatment-on-treated estimate was 0.28 standard deviations. That does not mean tutor quality is irrelevant. It means a strong system can help a less-experienced tutor succeed.
A mentoring relationship was not clearly driving those gains. Saga students did not report more or stronger adult relationships than the control group, despite improving academically. Relationships may matter for many reasons, but this study did not find evidence that they explained the test-score result.
A human need not fill every minute. In a later randomized study of roughly 4,000 students, four children shared a tutor. On a given day, two worked with the tutor while two used math software; the pairs alternated assignments every other day. The estimated gain among participating students was 0.23 standard deviations. The authors compared that with 0.28 in the earlier Saga trials and estimated that the hybrid cost about 30 percent less per student. These were separate experiments, not a head-to-head comparison. The defensible conclusion is simply that replacing some tutor time with structured practice preserved a sizable benefit in this program.
This is not a clean decomposition. But it suggests that attention matters, explanations can be supported by a strong curriculum, a mentoring relationship is not the only route to gains, and software can absorb some practice time. We still do not know how much each part contributes.
Modern Bloom will include children who start ahead, behind, and in the middle.
Students will be sampled randomly across two schools rather than selected by starting ability. The headline result will be the average effect of the full tutoring package. Modern Bloom has not published the pilot’s sample size, so we do not yet know whether it can reliably show if the effect differs by starting point.
Still, what happens to children at either end will be worth watching. Do those who start behind catch up? Do those already ahead keep accelerating? Individual trajectories cannot establish a separate effect, but they can show which questions a larger study should test directly.
Somers says the second-year design is not final, but will likely sample students across many schools. That would help show whether the result extends beyond one school or region.
Because everything changes at once, the comparison cannot separate level matching from mastery checks, extra practice, explanations, or personal attention. That is not a flaw. It is simply the question Modern Bloom has chosen: how far can the whole package move children under ideal conditions? Which ingredients mattered is a question for later research.
For parents, ask what the child does next - not what the product is called.
This question is personal for me. Last year, when Alex was in first grade, I taught him at home. Before the school year ended, he was already working on third-grade math and Chinese. This year we sent him back to public school. He is happy there. He has friends, enjoys his days, and his teacher reports that he is generous with helping other students who struggle with the material.
But academically, my sense is that he is coasting. Much of the second-grade material is work he already passed. I do not think public school is failing him. It is meeting one important need while leaving another unsolved. Our question is how to keep the community without making him wait a year for the curriculum to catch up.
So this is the checklist I’d recommend for parents:
Does the work stay at my child’s level? How is the starting point set, and can the work move above or below the enrolled grade as my child masters new material?
Is my child using the time well? Who notices idling, guessing, copying, or repeated mistakes?
Is the teaching good? Does it explain new ideas clearly, correct misconceptions, and give useful feedback rather than simply marking answers right or wrong?
Childhood is not a race through the curriculum. Friends, play, belonging, a love of learning, and family sanity all count. But a child enjoying school does not tell us whether the math lesson is useful, just as a fun app does not tell us whether the child is learning.
If Somers and Hastings show that children can move extraordinarily far under ideal conditions, that will be a genuine achievement. It will not mean every child needs a $100,000 adult. It will mean that an exceptional tutor made a year of good decisions: what to teach next, when to ask for another attempt, when to explain, and when to move on.
The next question is which decisions required the tutor - and which the rest of us can copy.
Sources and further reading
The study protocol has not yet been preregistered, so details beyond the official site and launch materials should be treated as provisional.
Benjamin Somers. Modern Bloom launch announcement.
Benjamin Somers. Email correspondence with the author, August 3, 2026.
Bloom, B. (1984). The 2 Sigma Problem.
Nickow, A., Oreopoulos, P. & Quan, V. The Impressive Effects of Tutoring on PreK-12 Learning.
Kraft, M. (2020). Interpreting Effect Sizes of Education Interventions.
Kraft, M., Schueler, B. & Falken, G. What Impacts Should We Expect from Tutoring at Scale?.
Duflo, E., Dupas, P. & Kremer, M. Peer Effects, Teacher Incentives, and the Impact of Tracking.
Banerjee et al. Mainstreaming an Effective Intervention: Evidence from Randomized Evaluations of Teaching at the Right Level in India.
Muralidharan, K., Singh, A. & Ganimian, A. Disrupting Education? Experimental Evidence on Technology-Aided Instruction in India.
de Barros, A. & Ganimian, A. Which Students Benefit from Computer-Based Individualized Instruction?.
Reis et al. Curriculum Compacting Study.
Kulik, J. & Kulik, C. Meta-Analytic Findings on Grouping Programs.
Steenbergen-Hu, S., Makel, M. & Olszewski-Kubilius, P. What One Hundred Years of Research Says About Ability-Grouping and Acceleration.
Pedersen, B., Makel, M., Rambo-Hernandez, K., Peters, S. & Plucker, J. (2023). Most Mathematics Classrooms Contain Wide-Ranging Achievement Levels (published version).
Koedinger et al. An Astonishing Regularity in Student Learning Rate, with Lee et al.’s 2026 reanalysis.
Robinson et al. The Effects of Virtual Tutoring on Young Readers, and Hsieh et al. The Power of Personalized Attention.
Guryan et al. Not Too Late: Improving Academic Outcomes Among Adolescents.
Bhatt et al. Can Technology Facilitate Scale?.
Paul von Hippel. Two-Sigma Tutoring: Separating Science Fiction from Science Fact.
Daisy Christodoulou. Bloom’s famous 2 sigma tutoring paper is incredibly misleading.








