Search papers, labs, and topics across Lattice.
This paper explores the inverse problem of identifying probabilistic structures from vanishing binomials in empirical probability tensors, utilizing algebraic signatures to facilitate structural learning. By focusing on the Kronecker-stack class of configuration matrices, the authors develop a method that matches signatures without the need for parameter estimation, thus streamlining the identification process. The approach was validated on both synthetic and real language data, revealing interpretable rank-one structures that correspond to meaningful sets of words, highlighting the practical applicability of algebraic statistics in computational linguistics.
Identifying probabilistic structures from vanishing binomials reveals interpretable language patterns without traditional parameter estimation.
Algebraic statistics characterizes statistical models through polynomial constraints, but it has mainly been used for analytically specified model classes. This paper studies the inverse problem: identifying probabilistic structure from vanishing binomials observed in empirical probability tensors. We treat the vanishing binomials of a toric model as its algebraic signature, and turn the ideal-variety correspondence of algebraic statistics into an operational procedure for structural learning that identifies a model by signature matching without parameter estimation. By restricting attention to a computationally tractable class of configuration matrices, which we call {\it the Kronecker-stack class}, we make these signatures explicitly enumerable. Within this class we define minimum invariant constraint (MIC) as the atomic unit characterizing each signature and generalizing the notion of independence. We tested this approach employing MICs on synthetic data as well as on corpus-scale real language data. The results suggested the utility of the method, revealing that the identified rank-one structures correspond to interpretable sets of words. These results open up a new avenue for applying algebraic statistics to computational linguistics.