Explainable AI in Secure Federated Learning: A Double-Edged Sword

As Federated Learning (FL) becomes a cornerstone for decentralized intelligence in 6G networks, its security vulnerabilities are coming under intense scrutiny. Because FL relies on aggregating model updates from distributed clients, a single malicious node can inject poisoned data or backdoors into the global model, compromising the entire network.

In Talk 04 of ENSURE-6G Event #6, researchers from the UCD School of Computer Science (NetsLab) explored a fascinating approach to this problem: using Explainable AI (XAI) to expose these attacks. However, as the presentation revealed, XAI is a double-edged sword—capable of both defending the network and arming attackers with highly stealthy evasion techniques.

The Threat: Poisoning Attacks in Federated Learning

In a standard FL setup, local clients train models on their private data and send updates to an aggregator. A malicious client can execute poisoning attacks with two primary goals:

  1. Degrading overall performance: Arbitrarily crashing the global model’s accuracy.
  2. Injecting backdoors: Manipulating the model to misclassify specific target data while behaving perfectly normally for everything else (making it incredibly hard to detect).

Traditional defenses often rely on distance-based algorithms (like Krum or FoolsGold) that try to spot anomalous weight updates. However, clever attackers can easily bypass these if they keep their weight deviations small.

The Defense: “Sherpa” and SHAP-based Detection

To combat stealthy poisoning, the UCD team developed Sherpa, a defense mechanism built on SHAP (SHapley Additive exPlanations).

SHAP is an Explainable AI technique that calculates the exact contribution (or “importance”) of each input feature to a model’s final prediction. The researchers discovered that even if a poisoned model and a benign model output the exact same prediction, how they arrived at that prediction—their feature attributions—will look vastly different.

  • Clustering for Detection: Sherpa analyzes the feature attributions of client models and groups them using a clustering algorithm (HDBSCAN).
  • Spotting the Anomaly: Benign models will cluster tightly together because they use the same logical features to make predictions. Poisoned models will exhibit highly deviant feature attributions, immediately exposing them.
  • Resilience: Because it analyzes logical behavior rather than mere mathematical distance, Sherpa successfully defends the network even when up to 80% of the clients are malicious—a threshold where traditional defenses entirely collapse.

The Dark Side: Using LRP to Craft Stealthy Attacks

While XAI provides a powerful magnifying glass for defenders, attackers can look through the exact same glass. The presentation demonstrated how adversaries can leverage another XAI technique—Layer-wise Relevance Propagation (LRP)—to craft devastatingly stealthy model-poisoning attacks.

LRP traces a prediction backward through a neural network, assigning “relevance scores” to individual neurons. It acts like a brain scan, highlighting exactly which neurons “fired” to reach a specific decision.

An attacker can use LRP to pinpoint the exact handful of neurons responsible for a target classification. Instead of heavily poisoning the dataset or drastically altering the model’s overall weights, the attacker only subtly tweaks those few critical neurons.

  • Evasion: Because the overall model weights barely change, distance-based defenses cannot detect the anomaly.
  • Decoy Tactics: The researchers showed that attackers can even deploy traditional data-poisoning as a “decoy” to distract the aggregator, allowing their highly targeted, LRP-crafted backdoor to slip through unnoticed.

Conclusion

Explainable AI is moving beyond a mere debugging tool; it is becoming a critical battleground for 6G network security. It offers network administrators unprecedented visibility into model behavior, allowing for highly effective poison detection. However, as attackers weaponize the same transparency tools to manipulate individual neural pathways, the ENSURE-6G community must develop robust, adaptive security frameworks that anticipate this double-edged reality.

Watch the Full Talk:

Previous Article

Predicting the View: Personalized Federated Learning for 6G Virtual Reality

Next Article

Securing the Edge: Towards Trustworthy and Privacy-Preserving Federated Learning

Write a Comment

Leave a Comment

Your email address will not be published. Required fields are marked *