Should you break a complex task into smaller subtasks, have specialized modules handle each piece, and stitch them together? Or should you hand the entire task to a single model, training it end-to-end without rigid intermediate steps? This debate has split machine learning for years. The advantage of the former is that each module can be debugged independently and swapped out at will; the advantage of the latter is optimization freedom. A transparent glass fails to hang on a hook because it disappears from the depth map. The robotics experiments below put a concrete score on both approaches. This article first looks at the scores for each side, explains why and where modularization fell short, examines when decomposing tasks remains the right choice, and concludes with three criteria to help you decide whether a pipeline stage should stay or go in your own system.
A 7-DoF robotic arm stands in front of a workbench. The gripper lifts a transparent glass from the tabletop and slowly moves it toward a hook. In front of the hook, the arm drifts off position, the handle misses the hook, and the glass fails to catch. The trial ends in failure.
To hang a glass reliably, a robotic arm typically relies on a dedicated depth-sensing hardware module, most commonly an RGB-D camera. Most of these cameras operate on infrared light: they project structured patterns onto objects, measure the pattern’s deformation as it reflects back, and compute the distance from each pixel to the lens.
Glass, however, is naturally transparent. The projected infrared light passes straight through the cup and refracts away along its curved edges. As a result, the sensor receives no matching reflections. In the resulting depth map, the region where the glass sits contains no valid readings, leaving an empty hole right where the object should be.
Once the front-end sensor loses critical data, the downstream action policy network is forced to guess from incomplete inputs, inevitably throwing off the gripper’s end-effector positioning. When using such cameras on glassware, the same failures recur: misjudging the cup wall, missing the handle, and losing track of the cup mid-air. This stems directly from optical ranging physics. Independent metrology measurements clearly mapped out these boundaries: flat glass sheets cause minimal disruption, but refractive volumes paired with dark backgrounds cause widespread reading failures. Technical support in Intel’s RealSense community noted that any material transparent in visible light is inevitably transparent to stereo-based depth algorithms as well; the only way to measure it reliably is to manually spray developer powder. Microsoft reached a similar conclusion when handling an Azure Kinect bug report, classifying depth collapse on semi-transparent objects as By Design: Everything is working as expected.
The flaw in this two-stage architecture lies in the intermediate sensor step: it insists on converting continuous light rays hitting the lens into a rigid distance grid. When transparent objects cause depth drops, only corrupted data reaches the action network, and the downstream policy has no way to propagate its failure back up to correct the front-end readings.
When dedicated depth-sensing hardware fails on glass, what happens if you strip away that step entirely and let a neural network inspect the raw footage from two lenses directly? Can it figure out depth on its own? A team from Northwestern University and Stanford University ran controlled comparisons on a real robotic arm (paper, project page). They built an experimental testbed centered around a 7-DoF robotic arm, mounting cameras on both the workbench edge and the arm’s wrist. Each camera actually consists of two standard optical lenses separated by 6.3 cm, capturing simultaneous views from slightly different angles. By default, these lenses can compute a depth map onboard. But this experiment deliberately ignored that computed depth, feeding the raw RGB images from the left and right lenses directly into the policy network instead.
For comparison, they also retained traditional approaches: back-projecting sensor-measured depths into 3D spatial coordinates, commonly known as point clouds. They designed five evaluation tasks: placing a banana into a box, slotting toast accurately into a toaster slot, and hanging a plastic cup, a stainless steel mug, and a glass cup onto a mug tree. Each task began with 200 demos collected via human teleoperation to train a diffusion policy network. During formal evaluation, object placements were randomized each time, and each task was evaluated across 20 trials on the real robot.
With identical robot hardware, identical demonstration data, and comparable model parameter scale, merely altering the data representation fed to the model produced an almost fourfold difference in success rate. The breakdown of results: feeding both left and right images into the network with cross-view attention fusion achieved 59%; feeding the same two images concatenated along the channel dimension without cross-view interaction dropped success to 45%; providing only a single-view RGB image yielded 43%; appending a hardware-computed depth map to that single RGB image dropped performance further to 41%; and the worst performer, feeding measured 3D points into a dedicated point cloud network, achieved a success rate of just 14%.
The paper is currently under academic conference peer review (submission page), the codebase has not been officially open-sourced, and no independent third-party replications exist yet. Cloud compute provider Lambda is listed among the author affiliations; the authors acknowledged their compute sponsorship in a social media thread (author acknowledgement thread), and Lambda published a technical overview on its corporate blog (Lambda blog post). Testing each method across five tasks with 20 trials each totals 100 trials on the real robot, a relatively modest sample size. When estimating confidence intervals under a binomial distribution, the bounds remain fairly wide; these numbers indicate a clear trajectory rather than an exact pinpoint value.
When robotic arms act in the physical world, traditional engineering assumes the front end must compute depth first. The system is split into two modules: the front end handles ranging and reconstructs 3D coordinates, while the back end consumes those coordinates to drive actions, stitching data across the interface. Fixed intermediate steps or hardware calculations that do not participate in backpropagation cut off gradients at that boundary. If front-end ranging drifts, the downstream policy network has no choice but to consume distorted data, while the front end never receives feedback from action failures.
Each input modality was evaluated on the real robot for 20 trials across five tasks, totaling 100 trials. Among them, the cross-view attention fusion scheme achieved a 59% success rate. The channel concatenation scheme dropped to 45%, single-view RGB achieved 43%, single-view RGB with a hardware depth map dropped to 41%, and the 3D point cloud approach scored just 14%. The point cloud method placed last because real-world depth measurements are riddled with missing holes and noise; converting them into spatial coordinates requires manual preprocessing like cropping, downsampling, and coordinate frame alignment. The authors pointed out that any minor flaw in this preprocessing pipeline easily collapses policy success to zero.
The gap between 45% and 59% illustrates the impact of representation architecture. Channel concatenation and cross-view attention share similar parameter counts and receive identical raw footage. In the 45% configuration, features extracted from the two lenses do not interact before reaching the decision layers. In the 59% network, backpropagation from action loss forces corresponding pixels between left and right images to establish automatic geometric correspondence, all without any manual depth annotations.
Subsequent ablation and probing experiments confirmed how this representation emerges. Replacing the right-eye image with a duplicate of the left-eye image caused performance to plunge immediately. When researchers analyzed the trained model’s internal features using linear probes, they could consistently extract the relative 3D distance between the gripper and the target object, whereas untrained weights revealed no such information. Without any hand-engineered 3D coordinate schema, the neural network developed an internal spatial representation tailored for manipulation entirely through end-to-end action optimization.
The success of an end-to-end stereo policy does not mean explicit geometric computation is obsolete. When input data is sufficiently clean, computing 3D coordinates first remains engineering-competitive. In evaluations on the OBSBench benchmark for manipulation tasks, simple point cloud policies often outperform pure RGB or RGB-D configurations. NVIDIA’s comparative study on point cloud robustness also demonstrated that coordinate-based models exhibit stronger resistance to illumination shifts and viewpoint jitter. When supplied with noiseless, ideal depth in simulation, control policies relying on spatial coordinates achieve success rates over 90% in most tasks. In the dextrous hand grasping experiment DextrAH-RGB, an end-to-end stereo policy tied overall with a depth-based modular pipeline; in a specific bin-packing test across 11 objects, the depth-based approach won on three, while the rest tied.
Direct stereo input also comes with practical constraints. Stereo cameras require precise hardware synchronization and calibration, and the physical baseline between lenses creates a zero-disparity blind zone when the gripper gets very close to objects. Feature matching across left and right views can still fail on transparent surfaces. In 20 trials of the glass-hanging task, the attention-fused stereo model succeeded only three times (3/20).
There is no need to treat end-to-end training and modular pipelines as dogmatic camps. Whether a component should be split off into an independent module depends on its error rate on real-world data and whether downstream modules can leverage redundant information to self-correct. When the operating environment is clean and sensors are stable, modular components are much easier to debug and swap out. But when faced with real-world noise like reflections and refraction, where dedicated ranging stages fail repeatedly, keeping gradients connected and letting the entire network learn together is often far more effective.
Years ago at an industrial IoT company, a project came along that traditional engineering wisdom dictated splitting into multiple sub-models, estimated to take months. Switching to a single model trained end-to-end produced a working prototype in two days. Hand-crafted sub-modules proved fragile on messy field data, whereas end-to-end training drove feature extraction directly from the end objective, bypassing brittle human-designed interfaces. This complements the robotics findings: when a front-end stage has questionable stability on real-world data, end-to-end training demonstrates clear efficiency gains; when a front-end module consistently outputs high-fidelity features, modular decomposition accelerates engineering iteration.
When evaluating system architecture, a few concrete questions help clarify the choice. Have you measured how often this module fails in production? When it outputs incomplete or distorted intermediate results, does the downstream module have redundant signals to fill in the missing information? Can the upstream and downstream modules be cleanly integrated from an engineering perspective? If the module is reliable, the downstream stage has room to self-correct, or organizational constraints prevent end-to-end joint training, keep the module decoupled. But if the front-end failure rate is high and the downstream cannot absorb the bias, you should strongly consider connecting the gradients end-to-end.
The same trade-off recurs far beyond robotics. Retrieval-augmented generation (RAG) splits long documents into fixed-length chunks and filters them via similarity thresholds, severing retrieval from reasoning: no matter how capable the downstream model is, it cannot recover critical context dropped during retrieval. Agent architectures that insert hardcoded state machines between planning and execution prevent real-world feedback from reaching the planning model. The historical shift in NLP and computer vision, which abandoned hand-crafted features in favor of deep networks learning representations directly, followed the same logic.
Like a robotic arm missing the hook before a glass cup, every un-differentiable component inserted into a pipeline shrinks the search space accessible to global optimization. Before deciding to isolate a specialized stage, measure its failure rate in the real world first, then verify whether downstream components can digest those errors on their own.