3 chapters in, this is pretty good so far!
Not the most beautiful writing style, and not everything makes sense (e.g. the AI permanently blacklists employees from every minimum-wage job just for being late a few times), but overall, it seems like the author thought things through carefully.
Some takeaways from this story:
Also, it's very impressive to me that Manna predicted in 2003 that massively parallel video game chips[1] running in the cloud would be key to computer vision.
More precisely, the story featured a stage where researchers wired together video game consoles, and then another stage where they made customized massively parallel chips.
a stage where researchers wired together video game consoles, and then another stage where they made customized massively parallel chips.
That was a retrodiction/contemporary approach, not really a prediction. 'Beowulf clusters' were very trendy and came up all the time on Slashdot c. 2003, especially the PS2 ones (which were followed by PS3 ones to make use of the custom Cell processor).
Darn, I asked Claude if it was impressive before commenting and it didn't mention this. Misaligned!
And of The Gig Economy, an under-appreciated literary masterpiece of our time.
I have increasingly had this sense of being teleoperated and some of my friends have mentioned it as well.
I agree with the sentiment that the "trades" are long-term not safe from automation. Not only are these jobs more likely to have verifiable outcomes (there is a right way to wire house, less so a right way to restructure an organization) but the people who do these jobs do not have the societal power for regulatory capture.
In the short term, I am curious what your model of teleoperation implementation looks like. I envision enormous social pushback both from potential contractors of trade services and, more importantly, the labor that would be teleoperated. This is nothing capitalism can't blow through, sure, but give me a model for, say, electricians:
I think the last scenario is the most likely - I think there's years worth of PR that teleoperation would need to compete in the general market (see: self driving cars). But this scenario also limits the reach of the changes to the world.
It’s not clear to me why the regulatory and permitting hurdles would be lower for the established company that starts using teleoperation than for a new company doing the same thing. Maybe the new company gets slightly faster scrutiny if they start doing illegal stuff? But if an established company’s permit filing rates go down without a clear explanation or they start replacing licensed labor with teleoperated labor on permitted projects, that’s still going to be noticed quickly.
That said, I do agree that consumer reputation effects mean that the biggest impacts happen when these practices become widespread across established companies. I don’t think that necessarily takes super long though; if there’s efficiency to be had, then the current wave of private equity buyouts of established businesses in the trades will simply accelerate. Funds will be able to buy existing companies at a premium to their current fundamentals, change the existing firms’ business practices to crank up their profits, and either keep the resulting revenue stream or flip the business for a quick profit.
Sorry, my distinction was unclear. The hurdle would be lower for an org with an internal team of teleoperators as opposed to a new business that would need to find more than one client for their teleoperations services. It's basically a client relationship issue that would be more viable internal to a large company.
In re: reputation, I think you're vastly overestimating the speed of capital optimization. The process you describe will take 5 - 10 years. This is especially true here because either:
There's then this binomial distribution that makes this harder than it seems.
How about inspections? That needs even less manual skill - just looking and checking - and it is something everyone needs. I think it is good canary for teleoperating tech readiness.
If you can do inspections then you can already do a similar level of teleoperation to the GPS driving example, where the computer is planning the behavior of a human who is already skilled at the task. “Do the thing” followed by “fix this issue”, “fix that issue”, etc.
It‘s not obvious to me that the economic incentives would lead to applying that capability for inspections before applying it for teleoperation itself, so I think inspections are not actually a reliable canary.
I suppose yes. It’s hard for me to see how anyone is going to decide to trust an inspection that can’t minimally identify the failures it detects (if only by saying “this photo shows a defect” or similar), though maybe I’m missing something.
For a skilled technician, pointing to a failure generally is going to be enough to identify a remedial task even if the system can’t guide them through a more granular process of doing that task yet, and even if that task is going to be broken down into sub-tasks by the technician. My thesis is that this level of inspection is already the core of a minimum viable product for teleoperation, or at least that the line between teleoperation and other assistive technology gets blurry in that zone. If you can do an inspection and identify what and/or where the inspection failures are (rather than just pass/fail on the whole) then that’s pretty trivially restructured into a task checklist that can verify task completion and add remedial tasks to the list in response to performance on earlier tasks. At least at some level of task granularity, that clearly becomes teleoperation, and my instinct is that the only principled line to draw is between a task system that includes a verification loop and one that does not.
If the thought is that teleoperation needs to be more refined than this to start significantly distorting labor markets, then I do agree with that. But I think in many cases that outcome is likely to be achieved via (potentially rapid) incremental improvement of an initially-rudimentary teleoperation system rather than by a step change from no teleoperation to introducing teleoperation. My prediction is that there will often not be any other single point in that evolution that, by itself, changes the dynamic more than merely getting adoption for routine machine verification.
To me, it's not totally a question of skill but of where the bottlenecks are. I can see routine inspections internal to a company being teleoperated but I do not see certifying inspections being teleoperated.
Routine inspections being teleoperated goes back to the limits of individual companies. Though they probably shouldn't, a lot of businesses operate on the fix-it-when-it's broken technique, so they're not doing non-mandated inspections.
Inspections for certifications or regulation either done by the government (will not implement AI) or by some liability-bearing compliance team. Liability seems to be something we currently mostly assign to humans right now. Also, there's a lot of incentive for these inspections to be challenged -- if I get a 'C' on my health inspection from a teleoperated "untrained" inspector, I am going to make a stink about that.
So, again, we are stuck at the idea of PR, incentives, and where bottlenecks land.
How does a regulator verify that no hallucination happened during an inspection? Who is going to bear responsibility for a mistake AI makes and the teleoperator can't notice because of lack of competence?
I was thinking about inspection that owner/buyer/client does privately, so no friction or wrong incentives, but other comments made me realise such inspection will not be readily visible.
People wearing glasses with built-in cameras, connected to today's strongest AIs could already do a lot with a bit of scaffolding.
Disagree. I don't think frontier LLMs can even beat classic pokemon games using only raw in-game images as input. And that's even after being willing to accept a level of latency that would be totally unworkable in a kitchen or on a worksite.
Making smart decisions about what to do in real physical environments using only a video stream is feels further away from training distribution than being able to play a relatively simple turn-based video game.
In terms of projecting trends based on recent advancements - I suspect evidence points in the oposite direction. I think we're much closer to you being able to avoid the need for teleoperation when getting app store approval (by giving Astra access to your browser and a bit of scaffolding) than we are to teleoperation becoming a feasible strategy for guiding someone on how to learn an unfimiliar cooking technique
I don't think frontier LLMs can even beat classic pokemon games using only raw in-game images as input.
That's out of date:
Fable 5 ... also needs less scaffolding: for example, previous Claude models struggled to play Pokémon FireRed even with harnesses that gave them additional helpful tools, but Fable 5 beat FireRed with a minimal, vision-only harness.
You are right about pokemon - my mistake.
I still stand by the claim that if you try and use fronteir models today to teleoperate a physical real world task today they will not be useful, and the claim that being able to take the human out-of-the-loop when it comes to app-store approvals is comming sooner than wanting to bring an AI into-the-loop for a human doing something in an uncontrolled physical environment. (Aside from using super high-level instructions that effectively skip the physical challenge like "get the laundry detergent from Aisle 7" or "clean the bathroom")
Can you explain in more detail? Like when you say use if for DIY repairs, can you give example of exactly what the loop looks like?
I find it useful in the sense of getting it to tell me "yes sand back all the way to the wood and then use a product like this one for the first coat <link>"
But that feels like it's a lot more the case of the LLM being useful thanks to it's computer use tools and a big knowledge base of how this stuff is normally done. (E.g. not that different to having someone write you up a set of instructions before you even begin the task)
It doesn't feel useful when it comes down to actually observing and responding to a real physical environment (like it won't be able to watch you sand and then notice you're using poor technique, it won't be able to notice just by watching a camera feed that the tin of paint is running out too fast and you should go get another one while waiting for this coat to dry, etc.)
can you give example of exactly what the loop looks like?
Something like: "X has stopped working, can you walk me through the process of fixing it, or determining that I should bring in a professional? I'm reasonably handy, have a good range of tools, and don't want to break the law". Then I do what it says, and if I can't do something I take a picture and ask what to do next. For example, I recently replaced my water heater's anode with this approach.
Would you be willing to share the transcript?
I think I'm still confused about whether it's giving you high level guidance (in a way not unlike you watching a youtube video about it), versus it's genuinely "in the driver's seat" when it comes to observing+responding to the environment?
Yeah, neither can they make coffee. I would actually expect teleoperation to be possible in mostly manual situations, where the only complex part is the expert knowledge(e.g. running a given biology/chemistry experiment in a lab may require basic lab skills and very advanced chemistry knowledge).
In fact, I woud predict that a large portion of articles in natural sciences ~5 years from now, will have most of its evidence created by AI operating a bunch of bachelours in a lab.
This kind of tech is framed negatively but would also have huge benefits.
Living in poor countries, you will see some truly remarkable plumbing and electrical "creativity". There is a permanent severe shortage of trained professionals.
It would be a boon to such societies if anyone with hands could open their LetAiJesusTakeTheWheel app and have the AI drive them to a safe and code-compliant result.
The villain in the 1930s sci-fi tv series 'Buck Rogers' operates 'robot mines' where workers are forced to wear "robot amnesia helmets" that remove all their autonomy. The description is basically Amazon's Jennifer, but not quite as demeaning. So we're there in many ways.
Back, smile, eye hand and brain are the things you use humans for in industry. Computer vision, robot arms, and LLMs are making progress in replacing all of them.
recently had the slightly disturbing experience of my hand muscles being stimulated through electrodes in my arms to control my hand movement. if we get a software intelligence explosion faster than a robotics one, humans might be very literally teleoperated by AI systems.
https://rentahuman.ai/ has been popular for a while. Is this what you're thinking of?
When I look at why I expect the world to change a lot in the next few years, and why other people expect slower changes, I think a big component is disagreement on the extent to which AI will affect non-computer work. Sure, programming has sped up massively with Claude Code etc, and models like Astra seem posed to make similar changes to work with spreadsheets and other common business tools, but what about work that doesn't include computers at all?
The classic picture of AIs doing things in the world is robots, but I think a more realistic picture of the near future is computers telling people what to do. Leaning into the way the world has become very scifi, we could call this "teleoperating" people. Many things that are hard for robots are very easy for people, there are strong economic reasons that push towards teleoperation, and this bypasses many legal and social limitations on what AI can do. We should expect this to lead to large and rapid changes in the physical world.
One of the most widespread examples today is driving. I put my destination into the GPS, and it tells me what to do. I handle the low-level physical motions and responding to the local circumstances; the GPS has a broader view of the world and handles the strategy.
When I think about why this happened much earlier than the huge amount of "teleoperation" I expect to see soon, a few factors. Driving is a major human activity, so it was worth making navigation software at a time when AI wasn't very good yet, even though this meant a ton of human hours going into building the system. It was also a place where the strategic component was a very strong fit for automation. You can memorize the map with enough work, but even then you won't have real-time street-by-street traffic information. AI solved this problem so well we don't even call it "AI" anymore. On the other hand, driving is a realtime control problem in an unconstrained environment where people die if you screw up and you can't even always safely stop. This makes it hard to automate, but also would make it impractical for an AI to guide non-drivers through the process. The only reason Uber etc have been able to commodify driving as they have is that so many people already know how to drive.
Thinking about where else we might see this, most AI use today looks a lot like management. You figure out what you want it to do, and describe in detail. It asks you some questions up front and others while it works. After some churning you get some a work product to assess. Maybe there's more back-and-forth, or maybe it's good as is. You set strategy and give context; the AI handles the implementation. Today's AI is normally only applied to the implementation to the extent that the task can happen fully within the computer. In cases when the AI can't physically, legally, or intellectually do something, the most efficient path to completing the task will often be the for the AI to handle strategy while delegating to a human to fill these gaps.
To illustrate what this delegation pattern can look like, let's look at how I recently got my Whistle Synth app into the Mac App Store.
At a high level, I set the strategy: "Can you walk me through the process of getting this into the Mac App Store?" But everything after that was either handled by the AI or delegated back to me. It handled included figuring out what tasks needed to be done, modifying the implementation to be compatible with the App Store restrictions, building the app, and giving me instructions. And then it delegated to me to record a demo video involving whistling (physical), register as a Mac Developer (legal), and clean up its App Store description (intellectual).
This was mostly pure instruction-following on my part: I was being teleoperated. Here's one example:
This was relatively mindless work for me. Just like being navigated through a city I don't expect to return to, I didn't bother trying to learn how this worked. I was loosely paying attention to make sure I wasn't doing anything dumb, but for future more capable systems I expect people to stop attending even that little.
Once it finished walking me through submission I had to wait a few days for review. It was accepted in the first round with no reviewer comments. This is a pretty big deal: App Store rules are notoriously complex, the reviewers very picky, and as a first-time amateur Mac developer there's no way I would have gotten this all right on the first attempt pre-AI.
Even though this was an almost entirely within-computers case, the important thing here is the pattern: by following AI instructions I did something that would have taken me a ton of work to learn how to do alone.
Note that in this case I was both doing the high level strategy ("put this in the app store") and filling in gaps for the AI (clicking a blue plus in App Store Connect). As "teleoperation" becomes more common I expect some of this, as people automate away parts of their jobs. Other times I expect it will look like, for example, a highly AI-pilled startup founder directing AIs that direct employees. A lot like gig workers "below the API" today. I expect early iterations of these jobs to be frustrating, with the AI not delegating well. Then, as AIs get sufficiently good at directing and anticipating, they'll be pretty mindless, for better or worse, as you stop needing to think for yourself at all.
What sort of jobs might switch to being teleoperation? The top candidates are any where the physical motions are relatively straightforward, timing is not critical, and people today are paid a lot for their knowledge and judgement. If you had an expert looking over your shoulder and telling you what to do, I expect most of you could do most of the work of an electrician. In fact, that's the bulk of how electricians learn their trade: through apprenticeship. Same goes for mechanics, healthcare technicians, inspectors, etc: they combine physical and intellectual components, where it's the knowledge that keeps a random person off the street from being able to do the job. People wearing glasses with built-in cameras, connected to today's strongest AIs could already do a lot with a bit of scaffolding.
To have a large impact, teleoperated workers wouldn't need to be able to do 100% of an existing job category. As long as the parts that can and can't be done this way can be easily separated, 90% could be done by teleoperated novices, while some of the former professionals spend their time on the remaining 10%. When I think about how these other jobs are likely to go, I expect we start with ones without regulatory barriers: HVAC techs (typically unlicensed) before electricians (licensed) before surgeons (licensed + heavily regulated + realtime + high stakes). [1]
So, teleoperation is probably very economically productive. Is it a good thing? I think mostly no, for several reasons. The big one is that I expect it to speed up the rate at which AI advances turn into additional AI advances. This shortens the time our society has to figure out what to do about these massive changes, and increases the risk that immature technology is rolled out widely. Rushed deployment is more likely to lead to disaster, and there are many ways this could go extremely wrong. And by "extremely wrong" I mean "AI kills everyone wrong". Creating minds smarter than ourselves is the most consequential thing humanity has ever done or will ever do, and we have to get it right.
Which is why I'm heartened to see a lot of support, including from the CEOs of Anthropic and OpenAI, for managing the pace at which these systems become increasingly capable. But even if we held constant at the capabilities of models publicly available today (let alone trained but not yet released) I think widespread teleoperation is still very likely. I expect this to be a massive disruption, one very difficult to integrate into our existing societal system.
The first issue is just that I expect these to be unpleasant jobs with low negotiating power. Since there are many tasks that almost anyone could do if expertly advised, and the employer can easily filter out the people who can't or won't, there's very little to keep wages or working conditions up. Then add in competition from laid-off knowledge workers, and I expect unprecedented unemployment.
So even if we can avoid the large risks of losing control of the future, falling into AI-enabled authoritarianism, facilitating bioattacks, etc, how we handle a world in which most people can't find work that pays them enough to live on will be an serious challenge. I expect this will require very large scale redistribution. [2] I'm not sure this happens by default, but I think it's achievable with significant effort. And as a very small fraction of spending in a vastly larger economy it would be a much easier sell.
[1] For a future post:
[2] Looking at what there is already, the US does less than most rich countries, but even here we have medicaid, EITC, CTC, WIC, SNAP, SSI, TANF, Section 8, LIHEAP. We spend maybe 3-5% of GDP on means-tested programs. Then ~7-10% of GDP goes to things like universal public education and medicare which aren't directed specifically at the poor but are still effectively redistributive. Internationally there's been some of this, but much less; until recently the US was spending maybe 0.04% of GDP on the kind of foreign aid (ex: PEPFAR) that is really about helping the world's poorest, and then the private sector (Gates etc) adding maybe 0.1% of GDP.
Comment via: facebook, lesswrong, mastodon, bluesky, X, substack