A tractor that drives itself has to see people, animals and obstacles in an open field. We built the perception platform that does — colour, thermal, stereo, lidar and radar together.

Safety
Sensor fusion
Radar
Thermal
Lidar
Safety first. Automating agricultural processes using autonomous vehicles and robots has a huge potential for reducing manual labor and optimizing yield. However, self-driven vehicles pose a major safety risk. Therefore, the SAFE project investigated technologies for maximizing the safety of both humans and animals, using multiple perception sensors and state-of-the-art object detection algorithms.
A perception platform was developed including color camera, thermal camera, stereo camera, lidar and radar. A variety of datasets were collected, including the popular and publicly available dataset: FieldSAFE.






The project
Safer Autonomous
Farming Equipment.
SAFE was a joint research collaboration between two agricultural machine manufacturers, a robotics consulting firm and two research institutions. It explored technologies for maximizing the safety of both humans and animals around autonomous farming vehicles, while minimizing the workload and supervision needed by farmers.
Early in their careers, Peter Christiansen and Mikkel Fly Kragh did their PhDs on safety for autonomous farming vehicles in the project. Their role was to explore how well a variety of sensor technologies — cameras, lidar and radar — could detect and avoid obstacles in unstructured agricultural environments.
They designed and built an advanced perception system, synchronized and calibrated all its sensors, and used it for extensive data collection at seven locations in Denmark. One of those datasets, FieldSAFE, has been downloaded more than 10,000 times.



Detection
Detection
algorithms.
Several state-of-the-art detection algorithms were investigated — both traditional computer vision and deep learning for object detection (bounding box predictions), and fully convolutional neural networks for semantic segmentation (pixel-wise predictions).

3D point clouds
Hand-crafted features or deep learning
For 3D point clouds, either generated from stereo vision or directly available from a multi-beam lidar, we also compared traditional methods based on hand-crafted features with methods based on deep learning.
The traditional approach looks at all points in the neighborhood of a single 3D point. Based on features such as linearity, planarity, scatteredness and height, each point is classified using a support vector machine (SVM). The deep learning approach is a fully convolutional neural network operating on 2D range images.
There is still debate on which representation of a 3D point cloud is best suited for deep learning. Some methods use multiple 2D views of a point cloud, some a hierarchical voxel-based representation, some graph neural networks, and some transformer-based models.


Sensor fusion
Combining sensors to reduce uncertainty
Sensor fusion, or multi-modal fusion, combines sensor data from different domains to increase robustness and confidence. Combining multiple sensors should result in reduced uncertainty compared to the performance of the individual sensors.
Below is an example of lidar and camera fusion. Using the extrinsic calibration between the sensors and the intrinsics of the camera, 3D points can be projected onto the corresponding 2D image. The result shows pixel-wise and point-wise classification using a conditional random field (CRF) for fusion. The white boxes mark the qualitative improvements from fusing 2D and 3D information, and from fusing sensor data across subsequent frames (time).


Obstacle mapping
From detections to a map
For an autonomous system to be safe, obstacle detection must be followed by obstacle avoidance. Among other steps, that means transforming detections from the local sensor frames to the vehicle frame, possibly followed by local or global mapping.
For this, we investigated occupancy grid mapping combined with probabilistic fusion of inverse sensor models. The detections from all sensors and algorithms were transformed to a 2D top-down view and fused across both space and time. IMU and RTK GPS sensors were used for global localization.
The slider shows a drone-recorded orthophoto of a grass field with and without the obstacle map overlaid. The map was generated from a single traversal along the periphery of the field.
More work
See the other cases, or read about the systems behind them.