Background
Intelligent team robots must jointly perform perception, planning, and control in dynamic environments with occlusion, limited field of view, and constrained communication resources. A single robot or a single viewpoint is often insufficient due to occlusion, small distant objects, blind spots, and semantic inconsistency, resulting in unstable object detection and scene understanding. Multi-robot collaborative perception can improve environmental understanding by exploiting complementary viewpoints from different agents, but it also introduces challenges in cross-view alignment, information redundancy, bandwidth limitation, temporal delay, and semantic conflict. As identified in the current multi-agent semantic occupancy pipeline, late fusion suffers from three key bottlenecks: small-object loss, boundary blurring, and semantic conflict. Therefore, this project aims to develop feature-level fusion and communication-aware collaboration mechanisms, enabling robot teams to selectively exchange high-value perception information under limited communication resources and improve global object detection and occupancy mapping quality.
Research Objectives
The objectives of this project are as follows. First, we aim to establish a multi-robot, multi-view collaborative object detection framework to improve detection stability and perception coverage under occlusion and limited field-of-view conditions. Second, we aim to build a multi-agent semantic occupancy prediction and global semantic map fusion mechanism, allowing local occupancy predictions from individual robots to be aligned to the world frame and integrated into a consistent global scene representation. Third, we will move from late fusion to feature-level collaborative perception to reduce small-object loss, boundary blurring, and semantic conflict. Fourth, we will develop adaptive transmission and bandwidth-aware alignment so that agents can selectively transmit high-value features under dynamic communication constraints. Fifth, we will introduce uncertainty modeling and delay-robust fusion to improve robustness under incomplete sensing, communication delay, and cross-agent inconsistency. Sixth, we will establish a dual-track validation pipeline with simulation and real-world data to evaluate perception quality, transmission efficiency, and sim-to-real generalization, while supporting system-level integration with digital twin, intelligent scheduling, and collaborative control modules.
Methods
The proposed technical approach consists of five core modules. First, we will construct a multi-robot and multi-view data pipeline, including Isaac Sim scenes, the CoPed dataset, multi-view synchronization, pose calibration, and real-world data collection and annotation. Second, we will develop collaborative object detection models that take primary and auxiliary robot imagery as input, extract consistent representations using a shared encoder, and perform cross-agent information exchange through multi-view feature fusion and GNN-based collaboration methods. Third, we will establish a semantic occupancy construction pipeline, where each agent predicts local semantic occupancy, aligns the prediction to the world frame through pose alignment, and fuses it into a global semantic occupancy map. Fourth, to address the bottlenecks of late fusion, we will design feature-level fusion, adaptive feature selection, and high-value feature transmission so that the system can preserve the most critical information for global perception under limited bandwidth. Fifth, we will incorporate bandwidth-aware alignment, uncertainty estimation, and delay-robust fusion to handle communication delay, temporal asynchrony across agents, semantic conflict, and low-confidence regions. The system will be evaluated using metrics such as perception quality, transmission efficiency, map consistency, small-object preservation, and sim-to-real generalization.
Innovation
The innovation of this project lies in integrating collaborative object detection and semantic occupancy construction into a communication-aware collaborative perception framework, rather than optimizing only a single task or a single sensing node. First, the project advances from conventional late fusion to feature-level collaborative perception, enabling agents to exchange high-value features at an earlier stage and thereby reduce small-object loss, boundary blurring, and semantic conflict. Second, the project introduces adaptive transmission, allowing the system to selectively transmit information according to scene complexity, occlusion level, feature importance, and communication conditions, instead of transmitting full images or complete feature maps by default. Third, the project combines bandwidth-aware alignment, uncertainty modeling, and delay-robust fusion, making collaborative perception more suitable for real multi-robot systems with limited bandwidth, sensing uncertainty, and communication delay. Fourth, the project jointly supports object-level detection and occupancy-level scene understanding, enabling integration with downstream planning, collaborative control, digital twin systems, and intelligent scheduling, with strong potential for system-level deployment.
Expected Outcomes
The project is expected to deliver a collaborative perception technology suite for multi-robot, multi-view, and communication-constrained environments. The expected outcomes include: 1. collaborative object detection models and multi-view feature fusion modules; a multi-agent semantic occupancy prediction and global semantic map fusion pipeline; feature-level adaptive transmission and bandwidth-aware fusion mechanisms; uncertainty modeling and delay-robust fusion strategies; a dual-track validation platform combining Isaac Sim, the CoPed dataset, and real-world data; and 6. a perception module that can be integrated with digital twin, intelligent scheduling, IoT, and collaborative control subprojects. In terms of impact, this project will improve the environmental understanding capability of team robots under occlusion, limited field of view, small-object scenarios, semantic conflict, and communication constraints. It will reduce the reliance on a single sensing node and support future applications in warehouses, factories, service robots, and intelligent environments.