Active Lighting Strategies for Consistent Part Detection in Bins

Structured light and wavelength choice matter more than algorithms for reliable bin picking.

Cover illustration for “Active Lighting Strategies for Consistent Part Detection in Bins”

Consistent part detection in bins depends less on which pose-estimation algorithm a system runs and more on which active lighting method feeds that algorithm its data. The failure point sits upstream, in how the sensor illuminates the bin, and downstream software can only get a usable point cloud to work with if that choice is right.

Why bin picking fails at the lighting layer

A bin-picking system succeeds or fails along a chain: a camera captures an image, that image becomes a 3D point cloud, software segments individual objects out of the cloud, a pose estimator figures out each object's orientation, and a planner decides how the gripper should approach it. Every step after the first depends on the data the first step handed it. When the point cloud is sparse, noisy, or full of holes, no segmentation routine or pose estimator downstream can rebuild the geometry that was never captured. So lighting needs attention before algorithm choice does, not after it.

Not every bin-picking task is forgiving about illumination method. Some applications can succeed with several different approaches, but others, usually ones involving shiny or tightly packed parts, will only work with one specific technique, and picking the wrong one means the system fails regardless of how good its software is. Human eyes tolerate ambient light without difficulty, but it is far too inconsistent for a 3D sensor to rely on. That's why dedicated illumination is the accepted baseline even for simpler passive-light setups, and it's why most industrial 3D bin-picking systems run active illumination of some kind: they project their own structured or coded light onto the scene to generate visible surface features and hold precision steady. Most pose-estimation algorithms in use were built and tuned against clean, well-formed point clouds. Feed them degraded data and their accuracy drops sharply, not gradually. That asymmetry, fragile software paired with an unreliable upstream input, is what makes the lighting layer the place where bin-picking deployments succeed or stall.

How reflective and variable surfaces corrupt point cloud data

Two forces do most of the damage before any algorithm gets a chance to run: the physical reflectivity of the parts themselves, and the inconsistency of the environment surrounding them. Metallic objects piled in a bin create specular reflections and inter-reflections, where light bounces from one part to another before it ever reaches the sensor. So the point cloud comes back sparse, noisy, and missing whole sections, and 6D pose estimation built on that data simply fails. Captured scans of polished metal parts such as cylinders or washers show this directly: point clouds with floating clusters of disconnected points and noise scattered with no relation to the actual object surface. A system trying to detect a washer under these conditions has no reliable geometry to detect it from.

Lighting conditions compound the problem, and a small shift in ambient light changes depth-map noise enough to produce incorrect object positions. Summer sunlight through a factory window can saturate a sensor one week, and winter's cold-toned LED fixtures can introduce uneven contrast the next, so a single bin-picking cell can see the same bin two different ways depending on the season or even the hour. One automotive supplier ran into exactly this: engineers reported that their bin-picking robot detected shiny bolts without issue in the morning, then failed on the same bolts in the afternoon once they'd picked up a thin film of oil. Light reflection caused the failure, not the robot and not its software.

Research on anti-reflective coatings confirms the mechanism directly. When highly reflective industrial parts were sprayed with an anti-reflective coating, their depth maps came back with substantially less noise and far fewer missing values compared to the same parts unsprayed. That single comparison makes the underlying point hard to dispute: changing the physical conditions of the surface, not the algorithm reading it, is what cleaned up the data. A large baseline distance between camera and projector adds a separate failure mode on top of this, since shadowing occurs wherever the projector's light and the camera's view don't overlap, carving holes into the scene. Small parts near bin corners and edges are already hard to localize, so they take the brunt of that lost detail. None of this is damage an algorithm can repair after the fact, because the information was never captured.

Structured light and fringe projection for part detection

Structured light and its common variant, fringe projection, deliver the finest spatial resolution among active lighting methods. That is why they anchor so many industrial 3D bin-picking systems. A projector casts one or more patterns of light across the object surface, a camera captures how those patterns deform against the object's geometry, and triangulation between camera and projector positions reconstructs a dense 3D point cloud. Blue LED illumination is a common choice for the projected pattern, because it produces high-contrast fringes and it resists ambient-light interference better than broad white light does.

If you capture more patterns per scan, you get a denser, cleaner, more complete cloud of usable 3D points. Each additional pattern also adds time to the capture cycle, and in high-volume bin picking, where throughput is often the binding constraint, that trade-off between data density and cycle time has to be made deliberately.

Structured light's other real limitation is motion. Strong ambient sunlight washes out projected patterns, and rapid movement during capture smears them. That is why scan-and-stop capture, where the robot halts briefly for each scan, has dominated structured-light deployments for years. That constraint is starting to loosen. Researchers mounted a customized one-shot active stereo camera on an OMRON Adept S650 robot arm and modified the projector so it stayed illuminated continuously instead of firing in sync with discrete stops, so the sensor could capture multiple images in quick succession while the arm kept moving. Strong, continuous illumination let every target object get a 3-millisecond exposure, short enough to avoid motion blur, so the scan-and-stop overhead that traditional structured light has needed just collapsed. Where older scanners needed dozens of individual stationary scans to cover a multi-compartment bin, continuous illumination during motion can build the same complete scene in real time.

Wavelength choice adds a further layer of decision-making, particularly for metal parts. Solomon's SolScan industrial 3D camera uses a green projector instead of standard white light to improve contrast when it scans shiny metallic parts, because metal surfaces reflect different wavelengths unevenly, and green light happens to produce better fringe contrast on these surfaces. No published study has run a direct head-to-head comparison of blue versus green projection across a shared set of industrial parts, so this remains a genuine judgment call for whoever is specifying a system.

Where laser triangulation fits

Laser triangulation, sometimes called sheet-of-light scanning, works differently: a single line of laser light projects onto the object, a camera positioned at a known, calibrated angle images where that line falls, and the displacement of the line encodes depth directly. It belongs to the same active-illumination family as structured light, but its projection geometry handles specular metal surfaces without relying on wavelength selection the way green-light fringe projection does.

That geometry gives laser triangulation real precision on metal surfaces, so it works well when specular parts dominate the bin. The trade-off is speed: laser triangulation scans slower than structured light and typically needs the scanner and the part to move relative to each other to build up a full surface, since a single laser line only captures one slice of geometry at a time. That requirement for relative motion fits naturally with conveyor-based or linear-motion setups, where parts already pass beneath a fixed scanner, but it becomes a real constraint in static overhead bin configurations, where structured light needs far less mechanical accommodation to do its job.

Practical guidance on 3D imaging methods for bin picking groups laser triangulation with structured light and time-of-flight as a family of options, and the right one depends on the part's material, the line's required speed, and the geometry being captured. Time-of-flight deserves a brief mention in that grouping: it captures depth at higher frame rates by measuring how long light takes to return from a surface, but its spatial resolution trails well behind structured light, which keeps it out of contention for tight-tolerance part localization even where its speed would otherwise be attractive. For system designers, laser triangulation earns a place on the short list specifically when parts are specular metal and the application's cycle-time budget can absorb its slower scan rate. But it isn't a general substitute for structured light across high-mix or high-speed bins, where structured light tends to win because it captures area faster.

Wavelength, ambient suppression, and interreflection handling across methods

Choosing a method family is only the first decision. How that method handles wavelength, ambient light rejection, and multi-bounce reflection decides whether it actually holds up once installed on a factory floor, and these choices interact directly with each other.

Blue laser projection suppresses ambient light by emitting at a narrow, specific wavelength that a bandpass filter at the camera can isolate cleanly, screening out most of the daylight and overhead-lighting variation that would otherwise corrupt the scan. Blue LED projection, used in some structured-light systems, works on a related principle, concentrating its output in a narrower spectral band than ordinary white light and pairing that with its own bandpass filtering, though it rejects ambient light less effectively than laser-based approaches do. For factories where daylight through windows or shifting overhead fixtures introduces real variation across a shift, narrow-wavelength suppression of this kind is the standard answer.

But suppression alone can't solve every problem that a pile of metal parts creates. The Zivid 3 XL250 builds interreflection artifact removal directly into its processing pipeline, an acknowledgment that even correct wavelength choice and solid ambient suppression leave behind artifacts from multi-bounce light ricocheting between piled metallic parts, artifacts that hardware physics alone won't clear up. Green wavelength projection, the approach Solomon uses, targets a different mechanism entirely: metal surfaces reflect green light differently than other wavelengths, and that differential reflectance sharpens fringe contrast in the 2D image before triangulation even happens, which cuts down the noise that interreflection would otherwise introduce into the final cloud.

These two approaches solve different problems, and treating them as interchangeable is a mistake. Blue laser excels at rejecting ambient light; green projection excels at improving contrast on metallic surfaces. A bin-picking line fighting sunlight incursion through factory windows should lean toward blue. A line fighting the surface finish of the parts themselves should lean toward green. Both decisions need to be made before the system goes live, not adjusted after the fact once the deployment reveals a mismatch.

Eye safety is a separate factor with real operational weight. A blue laser classified as eye-safe can run at full industrial power in spaces where people are present, without the extra enclosures or safety interlocks that high-power white or UV structured-light sources would otherwise require. On a factory floor where workers share space with the cell, that difference affects installation cost and floor layout, not just raw sensor performance.

Gripper-mounted and robot-eye-in-hand lighting configurations

Where the light source sits relative to the part changes the problem it has to solve, independent of which illumination method is in use. A fixed overhead light always strikes the bin from the same direction. Any part oriented away from that angle reflects poorly no matter how well-chosen the wavelength is. A light mounted on or near the gripper moves with the robot instead, holding a consistent angle between light source and part surface regardless of how that part happens to be sitting in the bin.

How a metallic reflective part appears to a sensor depends heavily on the camera's viewing direction and how light is distributed across the surface. Researchers addressed this directly by mounting a light source at a fixed position relative to the camera, attached directly to the robot's gripper, and pairing that setup with a data-driven convolutional neural network rather than a hard-coded model of bin geometry, because viewing direction and light distribution jointly determine how reflective automotive parts show up in the captured image.

Placing the sensor near the end effector also reduces a second problem: occlusion at bin edges and corners. A sensor that works from above and at close range meets shallower shadow angles than a fixed overhead system does, because the gripper can approach parts directly instead of viewing them obliquely across the bin's walls. Realizing that benefit takes planning rather than simply bolting a light to the gripper and letting it move. Work on next-best-view selection for eye-in-hand calibration addresses exactly this, treating sensor-position planning as a deliberate step for maximizing coverage and data quality rather than an incidental byproduct of where the arm happens to travel.

None of this comes free. Gripper-mounted illumination means the arm has to be in motion before any useful data becomes available, which adds mechanical complexity and ties the sensing cycle to the robot's own kinematics in a way fixed overhead sensors never have to deal with. Placement, in other words, is a second axis of the lighting decision, sitting alongside the method itself, and some surface conditions, especially reflective parts in awkward orientations, genuinely need the flexibility that placement provides.

Where algorithm-side compensation genuinely helps

The strongest case against this entire argument comes from computer vision research itself. A substantial share of recent bin-picking work aimed at reflective surfaces focuses on improving 6D pose estimation rather than improving the data feeding it, treating lighting failure as a condition the algorithm should learn to tolerate. That position carries real engineering weight. Active illumination control that adjusts in real time adds cycle-time cost and mechanical complexity, while a well-trained model can in principle generalize across a range of surface conditions without any physical change to the sensor rig.

There's a genuine domain gap backing this argument further. Industrial cameras often capture high-quality grayscale images, but foundational vision models are mostly trained on large-scale RGB datasets, which creates a mismatch between what the camera produces and what the model expects. Metallic parts add textureless surfaces, strong reflections, and dense multi-instance clutter on top of that mismatch, complicating recognition even further. Even a well-lit scene can need a preprocessing pipeline to bridge that gap, so active lighting alone was never sufficient on its own.

Still, there's a hard ceiling that compensation cannot push past. Picking fails when data quality falls short of what CAD matching requires because the calculated picking position is wrong. Most existing pose-estimation algorithms weren't built for point clouds with missing regions, scattered outliers, and floating artifact clusters, and their performance drops sharply the moment those conditions appear, regardless of how sophisticated the model behind them is. Parts sprayed matte produced depth maps with substantially less noise and far fewer missing values than the same parts unsprayed, captured under identical algorithmic conditions. The algorithm never changed between those two scans. The data did, and the result changed because of it.

Preprocessing pipelines and learned tolerance for messy input extend how far a given lighting setup can be pushed before it breaks down. They don't change which illumination method actually suits a given surface type, and they lower the cost of a sound lighting decision.

Sources

  1. Next-Best-View Selection for Robot Eye-in-Hand Calibration

    Provided the basis for the discussion of next-best-view selection as a deliberate planning step for eye-in-hand sensor positioning to maximize coverage and data quality.

  2. A Low-Cost, High-Speed, and Robust Bin Picking System for Factory Automation Enabled by a Non-Stop, Multi-View, and Active Vision Scheme

    Provided the details on mounting an active stereo camera on an OMRON Adept S650 arm with continuous illumination enabling 3-millisecond exposures during motion.

  3. BioDet: Boosting Industrial Object Detection with Image Preprocessing Strategies

    Provided the observation that industrial cameras often capture high-quality grayscale images while foundational vision models are trained on RGB datasets, creating a domain gap.

  4. Active Detection and Localization of Textureless Objects in Cluttered Environments

    Provided the basis for the discussion of mounting a light source at a fixed position relative to the camera on the robot gripper and using a convolutional neural network to handle how reflective automotive parts appear under varying viewing directions.

Svetlana Vorobyeva

Staff Writer, Manipulation and Hardware

Svetlana trained as a mechanical engineer in Saint Petersburg before completing a robotics PhD at TU Delft, where she researched compliant gripper design; she has been writing for technical and trade audiences since leaving academia in 2018. Her coverage at The Cycle Time centers on the hardware and algorithmic challenges behind generalist manipulation and the gap between benchmark performance and real-world reliability.