The Cycle Time

Grasp Planning Under Partial Occlusion in Cluttered Bins

Robots learn to grasp hidden parts by inferring geometry they cannot see.

Contributing Editor · · 12 min read
Cover illustration for “Grasp Planning Under Partial Occlusion in Cluttered Bins”
Perception for mixed SKUs and irregular parts · October 1, 2026 · 12 min read · 2,790 words

A single overhead camera looks down into a bin of metal brackets. One bracket sits near the top, but half its body disappears under a neighbor, and its far edge falls outside the camera's line of sight. The robot arm above it has to decide where to place two fingers or a suction cup on a surface it cannot fully see, against a shape it can only partly confirm, without knocking the whole pile into a new configuration. That decision is not a matter of sharper optics or a higher-resolution sensor. A single camera view captures only one side of any object, and clutter adds further occlusion on top of that structural limit. The geometry the robot needs to plan around is incomplete no matter how good the camera is. The robot has to infer what it cannot observe: the hidden geometry, the true pose, and the safest approach path.

That reframing carries weight because so much of the published research on bin-picking has been built and tested in conditions that soften the problem it claims to solve. Desai, Sen, and Deb, writing in Intelligent Service Robotics, point out that much existing bin-picking research still rests on settings with minimal clutter and occlusion, understating what real industrial scenes actually demand. A dense, disorganized bin hides far more geometry from the camera than a tidy test bin with a handful of well-separated parts. Warehouses and fulfillment centers do not sort parts before a robot arrives. They stack, tumble, and pile.

Textureless and reflective surfaces strip away the surface cues the robot needs to infer hidden shape. Algorithms that can partially infer hidden shape from surface cues, edges, shading gradients, and reflectance patterns lose those cues entirely on the kind of bare metal or matte plastic components common in manufacturing and logistics. Geometry is obscured twice over: once by the neighbor blocking the view, and once by the absence of texture that would otherwise hint at what lies beneath.

The consequences of getting this wrong are not confined to a single failed pick. A missed grasp forces the system into a replanning cycle, and each cycle adds latency, adds wear on the gripper and arm, and reduces the throughput that justifies automating the task. The rest of this piece works through the major technical strategies built to address that inference gap, and each one attacks a different piece of what the robot does not know.

Why standard grasp pipelines fail

A conventional detect-then-grasp pipeline assumes three things are settled before grasp planning starts: the object's full shape, its pose in space, and the behavior of everything around it once contact begins. None of those assumptions survive contact with a cluttered bin. The robot's unknowns break into three categories, hidden geometry, uncertain pose, and unpredictable neighbor response. Standard pipelines treat each one as solved, and that assumption is why they degrade once real clutter enters the picture.

The gap in hidden geometry is a training-data gap as much as a sensing gap. Grasp estimation algorithms trained on complete point clouds run into partial, occluded observations at deployment time and produce weaker candidates than their benchmark numbers suggest, because the input distribution they see in the bin does not match the input distribution they were trained on.

Pose estimation carries its own version of the same failure. Six-degree-of-freedom pose estimation for textureless, randomly stacked workpieces is unreliable when it has to work from a single viewpoint, and the solutions that do work well tend to require object-specific training, which limits how quickly a system can be redeployed on a new part. A pipeline tuned to recognize one bracket geometry does not generalize automatically to the next one on the line.

The third gap is the hardest to close analytically: what happens to everything else in the bin once the robot touches something. Any grasp or push action in dense clutter sets off multi-object contact, and the physics of that contact, friction, inertia, and accumulated shape error, compounds across a sequence of actions in ways that are difficult to predict in advance. Bejjani and colleagues identify this as an open challenge in their work on occlusion-aware search under conditions of partial observability. Because of that, grasp candidates cannot simply be ranked once and executed. They have to be scored for execution feasibility and collision risk, and the system has to reacquire sensor data and replan after every attempt. What looks from the outside like one discrete task, pick the object, is structurally a closed loop of observation, inference, and correction.

Reconstructing what is hidden: shape completion before grasp generation

One direct response to hidden geometry is to reconstruct the missing shape instead of treating partial observation as something to work around. Shape completion moves the point of intervention earlier in the pipeline: rather than generating grasp candidates from an incomplete point cloud, the system first estimates what the complete object looks like, turning a partial-information problem into one with approximate, but complete, information.

Diffusion models have become a central tool for this. They can perform category-level 3D shape completion from a single partial depth observation, producing a full reconstructed geometry before any grasp candidate gets generated at all. That is a meaningful architectural shift from earlier approaches that simply tolerated incomplete data and tried to grasp around the gaps, toward a system that actively infers past the missing regions before planning begins. The 2025 Bin-Picking Perception Challenge at ICCV pushed occlusion-resilient pose estimation forward using geometry-aware transformers, prior-guided correspondence learning, and targeted synthesis of hard training examples, and its results show shape-level reasoning moving from isolated research demonstrations into the benchmark mainstream that the field measures itself against.

Desai, Sen, and Deb demonstrate a concrete version of this completion-then-plan structure: a segmentation model trained on synthetic data feeds a 3D reconstruction pipeline, and by their own analysis, the combined system achieves a 90 percent grasp success rate on a standard industrial manipulator working in heavily occluded scenes of textureless components. That figure comes from a single study, but it demonstrates the practical payoff of resolving hidden geometry before grasp planning rather than during it.

The limitation has to be stated without softening it: generative completion can hallucinate. A model can produce a surface that looks entirely plausible and is physically wrong, and a grasp planned against that hallucinated surface can place the gripper directly into collision with a neighbor the model never actually saw. Confidence in the reconstruction is not the same as correctness. Compounding that risk is the sim-to-real gap: a model trained exclusively on synthetic data has no guarantee of transferring cleanly to real sensor data, and some amount of adaptation using real examples remains necessary. Self-training methods can approach the performance of models trained entirely on real data using only a minority fraction of real examples, but even that minority fraction costs time and labor to collect for every new industrial part that comes down the line. Shape completion buys a great deal of capability, but it buys it on faith in a model's guess about what it cannot see, and that is exactly the gap that active perception exists to close.

Moving the camera to reduce uncertainty: next-best-view planning as active inference

Next-best-view planning takes the camera off a fixed track and turns it into a decision-maker in its own right. Rather than gathering additional images from convenient or arbitrary angles, the system chooses each new viewpoint specifically because of how much it should reduce uncertainty about the object it is trying to grasp.

Breyer and colleagues built a closed-loop planner that illustrates this distinction cleanly. The system re-plans at high frequency, continuously folding new sensor measurements into a volumetric map of the scene, and it uses a task-driven stopping rule to decide when a candidate grasp has become stable enough to execute, rather than running a fixed sequence of explore-then-grasp steps the way earlier systems did. The stopping rule carries most of the practical value here: the system commits to a grasp only once that grasp's predicted configuration survives the integration of new observations, which cuts down on false commitments driven by a transient, partial view that later observations would have overturned.

Other systems extend the same active logic with added modalities. VISO-Grasp, from Shi and colleagues in 2025, pairs a large vision-language foundation model with a velocity-field-based next-best-view planner and an open-vocabulary detector, constructing an instance-centric map of spatial relationships in the scene that guides both view selection and a sequential grasping strategy when a direct grasp is not yet feasible. A 2026 semantic-aware framework from Kweon and Jeon folds geometric and semantic information together into next-best-view selection, and reports strong simulated success rates along with a perfect 10 out of 10 successful grasps in real-world testing, all without requiring prior knowledge of where objects were located in the scene. Lower-cost hardware variants have also emerged: mounting an RGB-D camera on the robot's own wrist, fusing multiple views into a pose buffer over time, and processing raw stereo footage into refined depth maps brings active perception within reach of setups that cannot justify a high-end fixed 3D sensing rig.

None of this comes free. Every additional viewpoint costs motion time, and in a bin where objects shift between observations, each replanning cycle risks chasing a target that has already moved rather than converging on a stable grasp. That trade-off between how much a system looks and how fast it acts resurfaces later, once completion, active viewing, and rearrangement are considered together.

Occlusion-aware search through clutter

Shape completion and active viewing both assume the target is at least partly visible somewhere in the scene. That assumption breaks down when a target object sits fully buried beneath its neighbors, invisible from every angle the camera can reach without first disturbing the pile. At that point, reconstruction has nothing to reconstruct from, and another viewpoint has nothing new to see. The robot has to maintain a belief about where the target most likely is and choose actions that reveal it as efficiently as possible, acting before it can see the thing it is looking for.

Bejjani and colleagues built a hybrid planner for exactly this situation. It combines a learned probability distribution over the target's likely pose, updated from a stream of partial observations, with a receding-horizon planner grounded in physics simulation and guided by a reinforcement-learned heuristic. That heuristic is trained specifically to act on observations that still contain occlusion, rather than waiting for a clean, unobstructed view that may never arrive. The receding-horizon structure matters here for a specific reason: by interleaving short bursts of planning with execution, and continuously correcting the internal simulator against real observations, the system avoids letting small modeling errors compound across a long sequence of actions, which is the same neighbor-uncertainty problem raised earlier in a different form.

RF-Grasp pushes past the boundary of what vision alone can do. It grasps objects that are fully occluded and undetectable by any camera, using RFID tags to localize the target inside an unknown, unstructured environment. For certain industrial settings, the most direct answer to complete occlusion is a different sensing modality.

The approach does not scale cleanly. Reasoning about how objects relate to and occlude one another in a densely packed scene gets combinatorially harder as the number of object categories and packing configurations grows, and that growth in complexity is an open problem in the field rather than one current systems have solved.

Rearranging the scene before grasping: when the right move is a push, not a reach

Sometimes the fastest way to a successful grasp is not to see more or infer more, but to change the scene itself. When a target sits buried under layers of other objects, pushing a neighbor out of the way can convert an infeasible grasp into a feasible one, and it does so without requiring a better sensor or a better geometric model.

Lachhiramka and colleagues, writing in Autonomous Robots in 2025, built a system where the manipulator learns a sequence of strategic pushes to rearrange a cluttered scene around a partially or fully occluded target. A network called GR-ConvNet predicts grasp points for the target directly, while a separate soft actor-critic model decides when to push instead, guided by a clutter map that turns the surrounding mess into a quantitative score the system can act on. OPG-Policy, from Ding, Zeng, Wan, and Cheng, reaches a similar capability through a different route: an amodal segmentation module predicts the parts of the target that are occluded, and a deep Q-learning motion critic estimates the expected reward of pushing versus grasping, with a coordinator selecting whichever action the model expects to pay off. The decision is learned from outcomes rather than fixed in advance.

Dengler and colleagues apply the same logic to a more constrained environment: confined shelf spaces, the kind found in warehouse racking rather than open bins. Their deep reinforcement learning planner decides, at each step, whether to reposition the camera for a new viewpoint or execute a push, using a 2.5D occupancy height map as its representation of the scene, and it constrains pushes to be minimally invasive so they do not disturb the layout more than necessary. That constraint reflects a real-world preference, not just a performance metric: a warehouse operator does not want a robot that solves its own uncertainty by scattering every other item on the shelf.

What runs through all three systems is not that robots are capable of pushing objects, which has been true for years. What is new is that these systems have learned when pushing beats grasping or viewing, as a decision made case by case rather than treated as a fallback used only when everything else fails. Hu and colleagues, applying equivariant models to sequence pushes and grasps together, report a 49 percent improvement in grasp success rate in simulation against strong baselines, with real-world testing showing gains of a similar scale. Which strategy wins in a given bin, completion, active viewing, or rearrangement, is set by the geometry of the clutter itself, not by a preference built into the system beforehand.

Foundation models and closed-loop perception redefine the planning boundary

The three strategies covered so far, completion, active viewing, and rearrangement, were developed largely as separate lines of research. They are converging now into single systems where perception is not a fixed step that happens before planning, but a component embedded inside the planning loop itself, continuously informing and being informed by it.

VISO-Grasp is a clear example of that convergence: it combines a large vision-language model with geometric next-best-view planning and an open-vocabulary detector to handle target-driven grasping under severe occlusion. The language model contributes a kind of spatial reasoning that pure geometric methods cannot produce on their own, which lets the system operate without a pre-built model of every object it might encounter. ActiveGrasp, from Lei and colleagues in 2026, selects views that most reduce the entropy of a calibrated grasp-pose distribution model, tightening the link between perception action and grasp outcome prediction into a single optimization objective.

This shift is not confined to academic publishing. A multi-jurisdictional patent filing strategy of that kind signals that a major industrial player views joint perception-and-planning, not perception accuracy taken alone, as the frontier worth defending competitively. Alongside that, benchmark infrastructure is catching up to the industrial reality the field has been slow to test against: work such as Efficiently Manipulating Clutter via Learning and Search-Based Reasoning, the XYZ-IBD dataset, and the 2025 Bin-Picking Perception Challenge at ICCV are built around severe occlusion, reflective surfaces, and tightly packed scenes, replacing the tidier academic test conditions that earlier benchmarks relied on and raising the evidentiary bar for any claim of real-world readiness.

The open problems that no current approach has resolved

The sim-to-real gap persists for any novel part a system has not been trained on, since synthetic training data alone does not guarantee transfer to a real sensor and a real gripper working against real friction. The scalability of relational reasoning in dense clutter remains unresolved, since the combinatorial space of how objects can occlude and support one another grows faster than current planners can search as object categories and packing densities increase. The trade-off between speed and accuracy sits unresolved between passive completion, which is fast but can hallucinate geometry it never observed, and active perception, which is more reliable but costs time the system may not have in a high-throughput setting.

None of that diminishes what has been built. It describes, with the precision the field itself uses, exactly where the boundary of current capability sits, and why grasp planning under partial occlusion remains, at its core, an unresolved inference problem rather than a solved engineering one.

Sources

  1. Robotic Grasping & Manipulation Planning 2026
  2. Occlusion-resilient pose estimation of textureless components in cluttered environment and its implementation in robotic bin-picking
  3. Viewpoint Push Planning for Mapping of Unknown Confined Spaces
  4. Closed-Loop Next-Best-View Planning for Target-Driven Grasping
  5. Occlusion-Aware Search for Object Retrieval in Clutter
  6. Learning Dual-Arm Push and Grasp Synergy in Dense Clutter
  7. VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF
  8. OPG-Policy: Occluded Push-Grasp Policy Learning with Amodal Segmentation

More in Perception for mixed SKUs and irregular parts