Loading prices…
STKR NewsSTKR News0 of 3 free this month
Regulation

Robots Are Learning to Feel

Computer vision isn't enough to make robots useful in the real world. New tactile datasets are finally giving machines the sense of touch they need to handle the messy reality of physical labor.

Originally on IEEE Robotics
AB

Adrian Boysel

Contributor

Sep 10, 2026

4 min read

Photo illustration / STKR News

We’ve been promised the age of robotics for a long time, but if you actually look at the current state of the art, it’s mostly just fancy cameras attached to clumsy metal arms. We have Vision-Language-Action (VLA) models that can identify a shirt or a kitchen gadget, but the moment a robot has to do something high-stakes—like turning a key in a rusted lock or handling a delicate egg—it usually fumbles. The reason is simple: robots are currently blind in their fingertips.

For founders and builders in the AI space, the current obsession is scale. We want more parameters, more GPUs, and more tokens. But in the physical world, scale isn't just about language; it’s about sensory input. Humans can do most complex tasks with their eyes closed because we have a feedback loop of force, friction, and slip. Robots don’t. We are currently trying to build the future of physical AI while ignoring one of the most important data streams in existence: touch.

The Vision Wall

In my view, we have hit a wall with pure vision-based robotics. You can feed a model every YouTube video of someone folding laundry, but that doesn't teach the model how much pressure to apply to a silk hem versus a denim seam. Vision is great for planning, but it’s terrible for execution. When a robot reaches for an object, the vision sensor is often obscured by the robot's own hand at the most critical moment—the point of contact.

Trevor Darrell at UC Berkeley is one of the few voices pointing out the obvious: you cannot simulate the nuance of a grip through a lens. His team recently demonstrated that adding just 100 hours of high-quality tactile data—covering basics like twisting and pouring—could double the success rate of complex tasks compared to the best vision-only models. That is a massive jump in performance for a relatively small amount of data. It tells me that tactile data is currently the highest-leverage asset in the robotics stack.

The Hardware Fragmenting Problem

The reason we don't have "GPT-4 for touch" yet is that tactile data is a mess. Unlike text or images, which have standardized formats (like .txt or .jpg), tactile sensors are all over the place. One robot might use a five-fingered hand with resistance sensors, while another uses a two-pronged pincer with gel-based cameras that track deformation. If you’re a developer trying to build a universal model, this hardware fragmentation is a nightmare.

However, we are seeing some clever workarounds. Researchers at Tsinghua University are trying to normalize this by mapping diverse sensor data onto a digital template of a human hand. They aggregated 3,000 hours of data from 21 different types of sensors. The result? A model that can "feel" even on hardware it’s never seen before. For builders, this is the blueprint. You don't need to own every robot on the market; you need a way to translate diverse physical signals into a common language.

Sparse Signals and Real-Time Stakes

There is a technical hurdle here that founders need to respect: latency. Vision models are relatively slow. They process frames and make decisions. But if a robot is about to drop a glass, it needs to adjust its grip in milliseconds. You can't wait for a heavy VLA to think about it. The Berkeley team solved this by using "expert" submodels—one for high-level planning (the brain) and a much faster one for low-level tactile adjustment (the reflex). This reflex-arc architecture is exactly how the human nervous system works, and it’s how we’ll get robots out of the lab and into warehouses.

The other issue is that tactile data is sparse. Most of the time, a robot isn't touching anything. If you feed a model a constant stream of "nothing," it learns to ignore the sensor. The Chinese Academy of Sciences is presenting a fix for this: a model that predicts what it should feel based on vision, and then only pays attention when the real tactile sensor tells it something different. It’s an anomaly detection approach to touch. If the robot expects to feel a hard surface but feels something soft, it amplifies that signal. This is a much more efficient way to train than just dumping raw data into a transformer.

The 100,000-Hour Milestone

So, where is the opportunity? Right now, the largest tactile datasets are around 30,000 hours. Compare that to the billions of hours of video data used for LLMs, and you see how early we are. Some experts, like Shunlin Lu at NeoteAI, think we need at least 100,000 hours of real-world tactile data to see a "Sora-level" breakthrough in robotics.

We are also seeing "tactile hallucination" techniques, where researchers at USC are training models to infer touch from video alone. By watching how a gripper interacts with an object, the AI can guess the pressure involved. This is a smart way to bootstrap, but I’m skeptical it can replace the real thing. You can't fake the physics of friction forever.

The Founder Takeaway

If you are building in AI, stop looking at just the screen. The next frontier isn't more chat bots; it’s physical agents that can actually interact with the world without breaking everything they touch. The bottleneck isn't compute anymore—it's high-fidelity, hardware-agnostic tactile data.

The Takeaway: Tactile intelligence is the missing link for "Physical AI." The winners in the next phase of robotics won't just have the best eyes; they'll have the best sense of touch. If you can solve the data normalization problem across different grippers, you’re sitting on a gold mine.


Read the original at IEEE Robotics →

The Brief

Stay Updated on Cutting-Edge Tech

A six-minute morning dispatch on the markets and the technology shaping them.

Free. No spam. Unsubscribe anytime.

Write for STKR

Become a Contributor

Earn $STKR for published stories on markets, protocols, and culture.

  • Earn $STKR for every published piece
  • Editorial support from the STKR desk
  • Byline visibility across the network
  • First look at the upcoming creator program
Apply to Write

Keep reading

All stories

Comments

24 reader responses