Aggregators

Creative Commons License

aGrUM

interactive online version

Aggregators are special type of nodes that includes a generic CPT for any numbers of parents.

pyAgrum proposes a list of such aggregators. Some of then are used below.

In [1]:
import pyagrum as gum
import pyagrum.lib.notebook as gnb
In [2]:
min_x = 0
max_x = 15

bn = gum.BayesNet()
l = [bn.add(gum.RangeVariable(item, item, min_x, max_x)) for item in ["a", "b", "c", "d", "e", "f"]]

gum.config["notebook", "histogram_line_threshold"] = 15
In [3]:
nmax = bn.addMAX(gum.RangeVariable("MAX", "MAX", min_x, max_x))
bn.addArc(l[0], nmax)
bn.addArc(l[1], nmax)
bn.addArc(l[2], nmax)
In [4]:
nmin = bn.addMIN(gum.RangeVariable("MIN", "MIN", min_x, max_x))
bn.addArc(l[3], nmin)
bn.addArc(l[4], nmin)
bn.addArc(l[5], nmin)
In [5]:
nampl = bn.addAMPLITUDE(gum.RangeVariable("DELTA", "DELTA", 0, max_x - min_x))
bn.addArc(nmax, nampl)
bn.addArc(nmin, nampl)
In [6]:
nmedian = bn.addMEDIAN(gum.RangeVariable("MEDIAN", "MEDIAN", min_x, max_x))
for n in [l[0], l[1], l[2], l[3]]:
  bn.addArc(n, nmedian)
# potential for median has a size : 16^5=2^20 double !
In [7]:
nexists = bn.addEXISTS(gum.LabelizedVariable("EXISTS_0", "EXISTS"), 0)
bn.addArc(l[0], nexists)
bn.addArc(l[1], nexists)
bn.addArc(l[2], nexists)
In [8]:
nforall = bn.addFORALL(gum.LabelizedVariable("FORALL_1", "FORALL"), 1)
bn.addArc(l[3], nforall)
bn.addArc(l[4], nforall)
bn.addArc(l[5], nforall)
In [9]:
ncount = bn.addCOUNT(gum.RangeVariable("COUNT_1", "COUNT_1,", 0, 3), 1)
bn.addArc(l[0], ncount)
bn.addArc(l[1], ncount)
bn.addArc(l[2], ncount)
In [10]:
for nod in l:
  bn.cpt(nod).fillWith(1).normalize()
In [11]:
gnb.showInference(bn, size="13")
../_images/notebooks_94-Tools_aggregators_13_0.svg
In [12]:
# dot | neato | fdp | sfdp | twopi | circo | osage | patchwork
gum.config["notebook", "graph_rankdir"] = "LR"
gnb.showInference(bn, size="13", evs={"MEDIAN": [0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 0, 0]})
gum.config.reset()
../_images/notebooks_94-Tools_aggregators_14_0.svg
In [13]:
# if the roots do not have uniform but random distribution
for nod in l:
  bn.generateCPT(nod)

gnb.showInference(bn, size="13")
../_images/notebooks_94-Tools_aggregators_15_0.svg
In [14]:
gnb.showInference(bn, size="13", evs={"MEDIAN": [0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 0, 0]})
../_images/notebooks_94-Tools_aggregators_16_0.svg

Input/Output

Aggregator (and ICI model) nodes have no stored CPT — the probability is computed on the fly from the aggregator’s rule. Saving to most BN file formats (BIF, BIFXML, DSL, XDSL, net, UAI, O3PRM) flattens that computed table into a full dense CPT before writing it out. Round-tripping through one of these formats therefore loses the aggregator: reloading the file gives back a plain node with a dense CPT, not an aggregator (bad round trip).

Only aGrUM’s native jgum (JSON) and bgum (binary) formats preserve the aggregator/ICI type exactly. Instead of a flat array of probabilities, the cpt entry for such a node is a small object describing its type and parameters, e.g. {"kind": "aggregator", "name": "forall", "value": 1} for a FORALL node, or {"kind": "ici", "name": "MultiDimNoisyORCompound", ...} for a Noisy-Or node.

In [15]:
import json
import os
import tempfile

demo = gum.BayesNet()
ps = [demo.add(gum.RangeVariable(f"P{i}", "", 0, 9)) for i in range(4)]
for i in range(4):
  demo.cpt(f"P{i}").randomCPT()
nsum = demo.addSUM(gum.RangeVariable("SUM", "", 0, 9))
for p in ps:
  demo.addArc(p, nsum)

with tempfile.TemporaryDirectory() as d:
  bif_path = os.path.join(d, "demo.bif")
  jgum_path = os.path.join(d, "demo.jgum")
  gum.saveBN(demo, bif_path)
  gum.saveBN(demo, jgum_path)
  bif_size = os.path.getsize(bif_path)
  jgum_size = os.path.getsize(jgum_path)
  with open(jgum_path) as f:
    jgum_content = json.load(f)

print(f"BIF  (dense CPT)         : {bif_size:>8} bytes")
print(f"jgum (compact aggregator): {jgum_size:>8} bytes")
print()
print(json.dumps(jgum_content, indent=2))
BIF  (dense CPT)         :   371370 bytes
jgum (compact aggregator):     1167 bytes

{
  "type": "BN",
  "GumJsonVersion": "1.0",
  "nodes": [
    "P0[10]",
    "P1[10]",
    "P2[10]",
    "P3[10]",
    "SUM[10]"
  ],
  "parents": {
    "P0": [],
    "P1": [],
    "P2": [],
    "P3": [],
    "SUM": [
      "P0",
      "P1",
      "P2",
      "P3"
    ]
  },
  "cpt": {
    "P0": [
      0.06566594173588765,
      0.08901794715652576,
      0.23549250176892222,
      0.29066539597222857,
      0.03538636695893471,
      0.021432408997347108,
      0.06830599300293316,
      0.07412138402999124,
      0.021497327645020436,
      0.09841473273220913
    ],
    "P1": [
      0.04956041578170093,
      0.052482099599744474,
      0.16902010720434282,
      0.06089639580981915,
      0.12238852929715743,
      0.052101213633520105,
      0.2224519701542751,
      0.013644502720908558,
      0.11585661283950999,
      0.14159815295902145
    ],
    "P2": [
      0.004132308655693356,
      0.32804681557541054,
      0.034186836296353496,
      0.01670458868887259,
      0.2148933374915759,
      0.13736848481534147,
      0.023250352497643734,
      0.0604824166875394,
      0.05557944492879496,
      0.12535541436277453
    ],
    "P3": [
      0.07497180006894322,
      0.3533814901330103,
      0.02646136631486612,
      0.03616975663656391,
      0.015781883892374837,
      0.05148092206412269,
      0.1696559152784467,
      0.05722052025284963,
      0.11232945089497481,
      0.10254689446384779
    ],
    "SUM": {
      "kind": "aggregator",
      "name": "sum"
    }
  },
  "properties": {
    "software": "aGrUM 3.2.0",
    "creation": "2026-09-28 15:47:45.972",
    "lastModification": "2026-09-28 15:47:45.986"
  }
}

Indeed, the cpt for the node SUM should have a size of \(10^4=10000\) floats. Instead :

In [16]:
print(json.dumps(jgum_content["cpt"]["SUM"], indent=2))
{
  "kind": "aggregator",
  "name": "sum"
}
In [ ]: