Search papers, labs, and topics across Lattice.
This paper explores the application of linear non-Gaussian acyclic models (LiNGAM) in federated learning environments, addressing the challenge of causal discovery while maintaining data privacy. The authors introduce the FedRCD family of algorithms, which leverage higher-order cumulant tensors to facilitate federated causal discovery without the need for extensive communication rounds, even in the presence of noise. Key findings reveal that these methods rank variables based on a variance ladder induced by the directed acyclic graph (DAG), rather than the population asymmetry typically used in centralised approaches, highlighting the limitations of existing methods like FedISHC under certain noise conditions.
Causal discovery in federated settings can achieve high accuracy without compromising privacy, even in the presence of noise, by leveraging higher-order cumulants.
In this paper we study linear non-Gaussian acyclic models (LiNGAM) when used in federated environments. These causal models allow one to go beyond Markov equivalence. However, in many domains data are scarce, and increasing the sample size by centralising data from different clients is not advisable due to regulations such as the GDPR. The federated environment offers an attractive option to balance privacy and causal discovery accuracy. Unfortunately, the standard centralised estimator in the LiNGAM setting, i.e., DirectLiNGAM, cannot be straightforwardly federated. Higher-order cumulant tensors offer a way around this obstacle: they depend only on the joint distribution of the variables involved and add exactly across independent sample groups, so a single communication round suffices in horizontal, vertical, and hybrid partitions. However, FedISHC, i.e., the current federated method along these lines, breaks down under near-symmetric noise. To overcome the above limitation, we introduce the FedRCD family of causal discovery algorithms, and investigate three variants that trade off communication rounds against algebraic noise; two of them are exact federated counterparts of the centralised high-order cumulant (HC) and HC-LiNGAM algorithms, and the single-round variants further effectively support exact unlearning at any granularity, from a single observation to a whole client. Numerical experiments show that at sample sizes typical of real deployments, the entire cumulant-based federated family does not actually rank variables by the population asymmetry that the scores encode at zero. It ranks them by a variance ladder induced by the DAG along its directed paths, the cumulant counterpart of varsortability. Marginal standardisation collapses every cumulant method to near-random ordering, while scale-invariant DirectLiNGAM, not federable under this protocol, is unaffected.