As 6G networks evolve, securing the underlying architecture against malicious activities is more critical—and complex—than ever. Traditional Network Intrusion Detection Systems (NIDS) often rely on centralized machine learning models, creating a bottleneck that poses both a single point of failure and significant privacy risks for distributed network nodes.
During Talk 01 of the recent ENSURE-6G Event #6, we explored a novel approach to resolving this bottleneck. Researchers from the University of Portsmouth and Telefonica presented their joint work on Decentralized Federated Random Forests for Network Intrusion Detection—a framework that delivers the robust accuracy of centralized machine learning while maintaining strict data privacy through a decentralized, peer-to-peer architecture.
(This research was recently presented at the prestigious ECML PKDD conference in Italy.)
The Privacy vs. Performance Dilemma in NIDS
Training a robust classification model to distinguish between normal traffic and anomalous, zero-day attacks traditionally requires massive amounts of tabular network traffic data.
- Centralized Learning: Aggregates all network traffic data to a central server. While highly accurate, it violates privacy constraints and introduces heavy communication overhead.
- Local Learning: Nodes train models solely on their own localized data. Unfortunately, this often results in unreliable detection due to highly heterogeneous data partitioning across different network domains.
- Centralized Federated Learning: Nodes train locally and only send model updates to a central aggregation server. This preserves data privacy but retains the central server as a potential point of failure.
To overcome these limitations, the presented research proposes removing the central aggregation server entirely.
The Solution: Peer-to-Peer, Decentralized Federated Learning
The framework leverages Random Forests (RF)—one of the most effective and lightweight ML algorithms for processing tabular network traffic data. However, instead of passing local RF decision trees up to a centralized coordinator, the model aggregates sequentially in a one-shot, peer-to-peer manner.
A local Random Forest is trained on the first node and propagated to the next. The second node filters and aggregates the best decision trees before passing the refined model onward. This process repeats until the global model is finalized at the last node and distributed back to all clients.
Key Findings and Performance Advantages
The framework was rigorously evaluated using solely numerical features across several benchmark NIDS datasets, including classic sets like KDD99 and UNSW-NB15, as well as more modern sets like CIC-IDS2017 and the 5G-NIDD dataset.
- Enhanced Accuracy: Decentralized Federated Random Forests actually outperform standard Centralized Federated approaches and achieve accuracy levels remarkably close to fully centralized models. By performing sequential aggregation, bad decision trees are filtered out much earlier in the peer-to-peer process, preventing them from diluting the global model.
- Immunity to Peer Ordering: The researchers tested sorting clients by training size, malicious traffic ratio, and purely random ordering. The results proved that the final model’s accuracy is independent of peer ordering. This eliminates the need for complex coordination protocols and ensures that no node needs to expose metadata about its local traffic to dictate the training sequence.
- Lightweight Inference: While training involves sequential propagation, the final Random Forest model deployed to network edge devices is highly efficient. Functioning essentially as a series of simple
if/elsestatements, the model requires minimal computational overhead—making it perfectly suited for resource-constrained 6G routers, switches, and Open RAN (O-RAN) components.
What’s Next for ENSURE-6G?
The transition from simulated benchmark environments to real-world deployment is the next critical phase. Future work will explore shifting from one-shot aggregation to iterative communication, evaluating adaptive tree selection under unreliable client participation, and exploring personalized federated learning setups.
Most importantly, the team plans to move this framework out of the lab. The next steps will involve deploying and measuring the training propagation times of this decentralized architecture directly on ENSURE-6G physical testbeds, pushing us closer to secure, privacy-preserving intelligence at the network edge.
Watch the Full Talk: