personal_asset

2026-09-12 Weekly Good Reads: AI Helping You Do Better Doesn't Mean You've Learned

A randomized trial of patent attorneys found that AI immediately increases output, but whether it consolidates into judgment independent of the tool depends on prior experience and how it is used.

2026-09-12 每周好文:AI 帮你做得更好,不等于学会

2026-09-12 Weekly Good Reads: AI Helps You Do Better, Not Necessarily Learn

Why I picked it

Over the past week, WeChat forwards kept circling around how ordinary people can carve out a future for themselves: is learning English still useful, what exactly are top salespeople doing, can education and personal effort change your circumstances. Behind all of them is really one shared question: is completing a task the same thing as gaining the ability to complete it independently later on.

A new paper by David Autor, Tanya Rodchenko, and others, Does AI Assistance Enhance or Erode Expertise?, drags this question into a real workplace. The research team partnered with 11 US intellectual property law firms and had 133 practicing patent attorneys take part in a three-month randomized trial. One group got a customized AI writing assistant, the other kept their existing tools; the study measured both how well they delivered while using AI, and whether they could still make professional judgments three months later when AI was taken away.

This paper is worth reading because it punctures two overly comfortable conclusions at once. The first is "AI lets newcomers quickly reach expert level, so experience no longer matters"; the second is "use AI and you'll degrade, so you'd better use it as little as possible." The results are messier, and closer to reality: junior attorneys gained the most and saved the most time on AI-assisted tasks, but showed no average improvement once AI was removed; senior attorneys had smaller immediate gains, but retained a clear boost in judgment after AI was taken away.

The paper's core argument

The study recruited 156 attorneys, of whom 133 completed enrollment and were randomly assigned. Randomization happened within each firm and was stratified by experience; 90 got Google Labs' not-yet-public InFlow writing assistant, 43 went into the control group. The paper defines seven-plus years of legal experience as senior, under seven as junior, and also ran robustness checks with thresholds ranging from three to nine years.

InFlow isn't just a chat box that continues your text. It can modify or insert content into selected text, generate text based on context, give feedback from the perspective of an editor, expert, or audience, and check for semantic inconsistencies according to rules. The experimental group also got a separate training session on patent writing, covering tool setup, prompt libraries, common errors and how to mitigate them; the control group only received general AI basics training and couldn't use generative writing tools.

The experiment set up three key assessments. About ten days after getting the tool, participants wrote dependent claims and detailed descriptions based on an invention's main claim, figures, and client notes; about ninety days later they did a harder patent drafting task. Finally, everyone had to revise a defective patent application without using AI — that is, redlining. The first two measured "how well can you do with the tool," and the last measured "does professional judgment stick around when the tool isn't there."

All work was blind-reviewed by patent attorneys at independent IP firms. Each piece was scored by two reviewers on enforceability, accuracy, ambiguity, completeness, and clarity. 105 people completed the ten-day task, 98 completed the ninety-day drafting, and 91 completed the final no-AI revision task; cumulative attrition from initial randomization to the last task was 32%, but the study found no significant differential attrition between the experimental and control groups.

First, immediate output. Attorneys using AI scored on average 0.34 standard deviations above the control group on the ten-day drafting task, and 0.38 standard deviations higher at ninety days, both statistically significant. The direction was consistent across all five quality dimensions, suggesting the gains weren't just from smoother prose.

Junior attorneys got bigger immediate help. At ten and ninety days, their estimated gains were 0.58 and 0.60 standard deviations respectively; senior attorneys were 0.22 and 0.30, with the former not significant and the latter only significant at the 10% level. On the ten-day task, junior attorneys in the experimental group were also about 18 minutes faster than junior attorneys in the control group; senior attorneys showed only about a 5-minute difference, indistinguishable from zero. AI really does lift newcomers' current performance and speed up their delivery.

Now, judgment after AI is removed. Three months later, the experimental group scored 0.32 standard deviations higher overall on the no-AI revision task. But once you break it down by experience level, this average result is entirely driven by senior attorneys: senior attorneys improved by 0.45 standard deviations, while junior attorneys' estimate was negative 0.03, essentially zero. In other words, the people who benefit most from AI for current output are not the ones who best convert the usage process into independent capability.

The junior attorney average also hides divergence. Junior attorneys in the experimental group had more dispersed no-AI scores: fewer mid-to-low pieces, but more bad pieces and more good pieces. The paper doesn't prove this as a definitive mechanism of "AI makes some people improve and others regress," only says the distribution is consistent with that interpretation. Senior attorneys' distribution shifted more clearly toward good work, with fewer bad and mediocre pieces.

The authors used text analysis to explore causes, but didn't treat it as a pre-registered causal test. Senior attorneys after three months of AI use more often restructured entire patent passages, tackled macro legal issues affecting commercial scope, and wrote legal reasoning for their revisions. Junior attorneys did more sequential formatting and local edits, could point out high-level problems, but often didn't actually complete the fix. The authors propose two possibilities: AI-generated drafts let experts shift attention to strategic judgment; for newcomers, the same offloading may crowd out the effortful practice needed to build strategic capability.

Points worth questioning

First, the sample is small and not representative of all knowledge workers. Only 91 people completed the final no-AI task, from 11 US law firms serving demanding clients in patent work. Patent writing is highly specialized, has strict confidentiality requirements, and tool use is controlled — the results can't be directly generalized to programmers, salespeople, students, or ordinary office work.

Second, the study didn't do a pre-enrollment no-AI ability test. The partner firms wouldn't accept that design, so the authors could only rely on baseline balance after randomization and interpret endline differences as experimental effects. Randomization is strong evidence, but directly comparing each person's before-and-after ability would have been clearer.

Third, cumulative attrition reached 32%. The paper's statistical tests found no differential attrition between experimental and control groups, and main results underwent robustness analysis, but in a small sample "not significant" doesn't mean there's absolutely no selective dropout. Estimates for the senior and junior subgroups especially need to retain uncertainty.

Fourth, the so-called "no AI" mainly relied on explicit instructions; the research team couldn't fully monitor whether participants used other tools outside the task. The authors checked InFlow logs, submission characteristics, and ran exclusion analyses, finding no evidence sufficient to overturn the results; but it still wasn't a physically isolated closed-book exam.

Fifth, the experiment only lasted three months, while professional capability typically takes years to form. Junior attorneys' divergence might converge later, or it might widen; whether senior attorneys' gains persist long-term also has no answer. The paper captured an important medium-term signal, not a career conclusion.

Sixth, the tool underwent feature changes during the trial, and direct costs were funded by Google, with multiple authors being Google employees or contractors who disclosed conflicts of interest. Work was blind-reviewed by external patent attorneys and the study was pre-registered — these arrangements reduce bias risk, but readers should still factor funding and product background into their judgment.

Seventh, the learning mechanism the authors propose is still an explanation, not a randomly manipulated variable. Senior attorneys didn't save obvious time yet accumulated more independent judgment, while junior attorneys saved more time but didn't accumulate on average — this relates to "whether you continue investing cognitive effort"; but the experiment didn't randomly assign thinking time, nor directly measure where each person spent the time they saved.

Takeaways tied to recent interests

This paper adds a crucial check to "is learning a skill still useful": don't just look at what you can deliver while using the tool, but also at what you can judge once the tool is taken away. Output gains are certainly valuable, but they don't automatically prove capability growth.

For people just starting out, the most dangerous thing isn't using AI, but outsourcing the very steps where you most need to form judgment. You can let AI give you drafts, explanations, and feedback, but keep at least three things for yourself: first judge independently where the problem is, then decide which suggestion to adopt, and finally redo it once without looking at the answer. If every time you just accept a result that looks professional, what you might be getting good at is the invocation process, not the underlying capability.

For people who already have experience, AI is more like an attention amplifier. Your existing domain model helps you identify which generated content is only superficially fluent and which parts touch real risk; once the tool generates the basic text for you, attention can shift to structure, boundaries, and trade-offs. "Experience" here isn't tenure itself, but the ability to explain why you're changing something, what the change will affect, and what can't be left to the model's judgment.

This also reminds us to rethink top salespeople and English learning. Memorizing a passage or completing one conversation is not the same level of capability as independently identifying the other person's real concerns or choosing the right expression. AI can spar with you, polish your writing, and simulate clients, but you should regularly add unassisted tests: can you grasp the problem without prompts, can you organize your views without a translation, can you handle unexpected follow-ups without a script.

For managers, you can't evaluate AI training only by "how much current output improved." At minimum, split it into two parts: quality and time with the tool present, and judgment, error-correction, and transfer with the tool absent. If newcomers only take on tasks simplified by AI, short-term numbers may look great, but the effortful yet skill-building steps of traditional apprenticeship will disappear.

The more realistic approach isn't banning AI, but redesigning work: let AI handle checkable drafts, require newcomers to explain key trade-offs and complete independent reviews; have senior people make their judgment process during review explicit, not just give final edits. Tools can speed up delivery, but organizations still need to preserve friction for capability accumulation.

How to read it

Start with the abstract and Figure 1, and note the three sets of results separately: ten-day AI-assisted drafting, ninety-day AI-assisted drafting, and ninety-day no-AI revision. Don't directly call the output gains in the first two "learning."

Then read Section 3 on experimental design, and check how the 133 people were grouped, what each of the three tasks measured, and how the completed sample dropped from 105 to 91. Then read Tables 3, 4, and 6, focusing on comparing point estimates, significance, and confidence intervals between junior and senior attorneys, rather than just remembering "AI works."

Finally read the limitations in the discussion: no baseline ability test, small and unusual sample, three-month observation period, inability to fully monitor the no-AI task. After reading, run a double acceptance check on one of your own AI workflows: is it faster and better with the tool present, and a week later can you independently spot the same kind of problem without looking at the original answer. Only when both improve is it closer to "capability growth."

NBER research page: Autor et al.: Does AI Assistance Enhance or Erode Expertise?

Full paper: NBER Working Paper 35720 PDF

Sources