aGrUM 3.2.0
a C++ library for (probabilistic) graphical models
gum::learning::KTBNDatabaseGenerator< GUM_SCALAR > Class Template Reference

Generates a database of trajectories from a k-DBN (one CSV per trajectory). More...

#include <agrum/KTBN/database/KTBNDatabaseGenerator.h>

Inheritance diagram for gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >:
[legend]
Collaboration diagram for gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >:
[legend]

Classes

struct  ParentRef
 a parent of a template node, precompiled for fast sampling More...
struct  NodeRef
 a template node, precompiled for fast sampling More...

Public Types

enum class  DiscretizedLabelMode : char { INTERVAL , MEDIAN , RANDOM }
 rendering of discretized variables when labels are requested More...
enum class  VarOrderMode : char { RANDOM , TOPOLOGICAL , ANTI_TOPOLOGICAL }
 column order used for the exported CSV More...

Public Member Functions

Constructors / Destructors
 KTBNDatabaseGenerator (const KTBN< GUM_SCALAR > &kdbn)
 Constructor.
 ~KTBNDatabaseGenerator ()
 destructor
Accessors / Modifiers
std::vector< double > drawSamples (Size nbSamples, Size nbTimeSlices, std::string_view dirPath, std::string_view csvBaseName, VarOrderMode mode=VarOrderMode::RANDOM, bool useLabels=true, std::string csvSeparator=",")
 Generates nbSamples independent trajectories, writing one CSV file per trajectory into dirPath.
std::vector< double > drawSamples (const std::vector< Size > &nbTimeSlices, std::string_view dirPath, std::string_view csvBaseName, VarOrderMode mode=VarOrderMode::RANDOM, bool useLabels=true, std::string csvSeparator=",")
 Like drawSamples(), but every trajectory may have its own horizon.
void setDiscretizedLabelModeRandom ()
 set discretized-label rendering to a uniform random draw in the interval (this is the default; each labelled export then differs)
void setDiscretizedLabelModeMedian ()
 set discretized-label rendering to the (deterministic) interval median
void setDiscretizedLabelModeInterval ()
 set discretized-label rendering to the interval label "[min,max["
Size nbVars () const
 returns the number of base variable columns

Public Attributes

Signaler< Size, double > onProgress
 Progression (percent) and time.
Signaler< std::string_view > onStop
 with a possible explanation for stopping

Private Member Functions

void _build_ (const KTBN< GUM_SCALAR > &kdbn)
 one-shot initialisation called by the constructor: fills the column index (baseCols, nbVars), the topological node/parent cache (nodes, kernel), the per-column representatives (vars), and the shared instantiation (inst). All cached pointers refer to template.
Idx _drawVar_ (const DiscreteVariable &var, const Tensor< GUM_SCALAR > &cpt, double &log2likelihood)
 inverse-CDF draw of var given the parents already set in inst; accumulates log2(P(drawn value)) into log2likelihood
std::string _label_ (Idx col, Idx idx) const
 renders the label of modality idx of base column col (taking the discretized-label mode into account)
void _writeTrajectory_ (std::string_view csvFileURL, const std::vector< Idx > &traj, Size nbTimeSlices, bool useLabels, const std::string &csvSeparator, const std::vector< Idx > &colOrder) const
 writes one trajectory CSV (header + T rows) to csvFileURL. traj is the flat row-major buffer (T x nbVars) in canonical column order; colOrder gives the output column order.
void setVarOrderRandomized (std::vector< Idx > &colOrder) const
 builds a uniformly random column order
void setVarOrderTopological (std::vector< Idx > &colOrder) const
 builds a topological column order (transition-kernel projection)
void setVarOrderAntiTopological (std::vector< Idx > &colOrder) const
 builds the reverse of setVarOrderTopological()
std::vector< double > _drawSamples_ (Size nbSamples, Size fixedLen, const std::vector< Size > *perTraj, std::string_view dirPath, std::string_view csvBaseName, VarOrderMode mode, bool useLabels, const std::string &csvSeparator)
 the single worker behind both drawSamples() overloads. Trajectory i's horizon is read from perTraj (when non-null) else from fixedLen; samples each trajectory and writes it straight to its own CSV file.
 KTBNDatabaseGenerator (const KTBNDatabaseGenerator &)=delete
 KTBNDatabaseGenerator (KTBNDatabaseGenerator &&)=delete
KTBNDatabaseGenerator & operator= (const KTBNDatabaseGenerator &)=delete
KTBNDatabaseGenerator & operator= (KTBNDatabaseGenerator &&)=delete

Static Private Member Functions

static std::pair< std::string, int > _decode_ (const std::string &name, const std::unordered_set< std::string > &temporalSet)
 decodes a template node name into (base, slice): "B[t]" with B a known temporal process -> (B, t); anything else (bare or bracket-named atemporal node) -> (name, ATEMPORAL).

Private Attributes

BayesNet< GUM_SCALAR > _template_
 the \(k\)-slice template (a small copy, independent of the horizon)
Size _k_
 the order \(k\) of the k-DBN
Size _nbVars_
 number of base variable columns
std::vector< std::string > _baseCols_
 col index -> base name (canonical column numbering)
std::vector< const DiscreteVariable * > _vars_
 one representative variable per base column (same order as baseCols), pointing into template so it outlives the source k-DBN. Label rendering only.
std::vector< NodeRef > _nodes_
 all template nodes in topological order (drives Phase 1, the bootstrap)
std::vector< Idx > _kernel_
 indices, in nodes, of the slice-(k-1) nodes (the transition kernel, drives Phase 2); already in topological order
Instantiation _inst_
 a shared instantiation over all template variables, so we don't have to rebuild it for every draw
DiscretizedLabelMode _discretizedLabelMode_ = DiscretizedLabelMode::RANDOM
 rendering of discretized variables when labels are requested

Detailed Description

template<GUM_Numeric GUM_SCALAR>
class gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >

Generates a database of trajectories from a k-DBN (one CSV per trajectory).

See also
gum::learning::BNDatabaseGenerator, the (static) Bayesian-network counterpart. Same role and sampling principle, but it streams trajectories from the k-DBN template without unrolling and without keeping the whole database in memory.

Definition at line 108 of file KTBNDatabaseGenerator.h.

Member Enumeration Documentation

◆ DiscretizedLabelMode

template<GUM_Numeric GUM_SCALAR>
enum class gum::learning::KTBNDatabaseGenerator::DiscretizedLabelMode : char
strong

rendering of discretized variables when labels are requested

Enumerator
INTERVAL 
MEDIAN 
RANDOM 

Definition at line 111 of file KTBNDatabaseGenerator.h.

111: char { INTERVAL, MEDIAN, RANDOM };

◆ VarOrderMode

template<GUM_Numeric GUM_SCALAR>
enum class gum::learning::KTBNDatabaseGenerator::VarOrderMode : char
strong

column order used for the exported CSV

Enumerator
RANDOM 
TOPOLOGICAL 
ANTI_TOPOLOGICAL 

Definition at line 114 of file KTBNDatabaseGenerator.h.

114: char { RANDOM, TOPOLOGICAL, ANTI_TOPOLOGICAL };

Constructor & Destructor Documentation

◆ KTBNDatabaseGenerator() [1/3]

template<GUM_Numeric GUM_SCALAR>
gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::KTBNDatabaseGenerator ( const KTBN< GUM_SCALAR > & kdbn)
explicit

Constructor.

Parameters
kdbnThe k-DBN to sample from (only its \(k\)-slice template is copied, via toBN(); the k-DBN itself is not retained). The horizon is not fixed here; it is passed to drawSamples().

Definition at line 77 of file KTBNDatabaseGenerator_tpl.h.

77 :
78 _template_(kdbn.toBN()), _k_(kdbn.k()) {
81 }
Generates a database of trajectories from a k-DBN (one CSV per trajectory).
void _build_(const KTBN< GUM_SCALAR > &kdbn)
one-shot initialisation called by the constructor: fills the column index (baseCols,...
KTBNDatabaseGenerator(const KTBN< GUM_SCALAR > &kdbn)
Constructor.
BayesNet< GUM_SCALAR > _template_
the -slice template (a small copy, independent of the horizon)

References KTBNDatabaseGenerator(), _build_(), _k_, and _template_.

Referenced by KTBNDatabaseGenerator(), KTBNDatabaseGenerator(), KTBNDatabaseGenerator(), ~KTBNDatabaseGenerator(), operator=(), and operator=().

Here is the call graph for this function:
Here is the caller graph for this function:

◆ ~KTBNDatabaseGenerator()

template<GUM_Numeric GUM_SCALAR>
gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::~KTBNDatabaseGenerator ( )

destructor

Definition at line 84 of file KTBNDatabaseGenerator_tpl.h.

References KTBNDatabaseGenerator().

Here is the call graph for this function:

◆ KTBNDatabaseGenerator() [2/3]

template<GUM_Numeric GUM_SCALAR>
gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::KTBNDatabaseGenerator ( const KTBNDatabaseGenerator< GUM_SCALAR > & )
privatedelete

References KTBNDatabaseGenerator().

Here is the call graph for this function:

◆ KTBNDatabaseGenerator() [3/3]

template<GUM_Numeric GUM_SCALAR>
gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::KTBNDatabaseGenerator ( KTBNDatabaseGenerator< GUM_SCALAR > && )
privatedelete

References KTBNDatabaseGenerator().

Here is the call graph for this function:

Member Function Documentation

◆ _build_()

template<GUM_Numeric GUM_SCALAR>
void gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_build_ ( const KTBN< GUM_SCALAR > & kdbn)
private

one-shot initialisation called by the constructor: fills the column index (baseCols, nbVars), the topological node/parent cache (nodes, kernel), the per-column representatives (vars), and the shared instantiation (inst). All cached pointers refer to template.

Definition at line 108 of file KTBNDatabaseGenerator_tpl.h.

108 {
109 const std::unordered_set< std::string >& temporalSet = kdbn.temporalVarNames();
110
111 // 1) canonical column numbering: one column per base variable (temporal
112 // variables first, then atemporal ones).
113 std::unordered_map< std::string, Idx > colByName; // local: only needed during construction
114 _nbVars_ = 0;
115 for (const std::string& base: temporalSet) {
116 _baseCols_.push_back(base);
117 colByName.emplace(base, _nbVars_++);
118 }
119 for (const std::string& base: kdbn.atemporalVarNames()) {
120 _baseCols_.push_back(base);
121 colByName.emplace(base, _nbVars_++);
122 }
123
124 // 2) precompile every template node (and its parents) in topological order,
125 // so that drawSamples() never has to parse a name again. In the same pass we
126 // pick one representative variable per column (_vars_, for label rendering)
127 // and build the shared instantiation (_inst_, reused for every draw). All
128 // cached pointers refer to _template_ (our own copy), so they outlive the
129 // source k-DBN.
130 _vars_.resize(_nbVars_);
131 const int lastSlice = int(_k_) - 1;
132 for (const NodeId n: _template_.topologicalOrder()) {
133 const DiscreteVariable& v = _template_.variable(n);
134 const Tensor< GUM_SCALAR >& cpt = _template_.cpt(n);
135 const auto [base, slice] = _decode_(v.name(), temporalSet);
136
137 NodeRef nr;
138 nr.var = &v;
139 nr.cpt = &cpt;
140 nr.slice = slice;
141 nr.col = colByName[base];
142
143 for (Idx i = 1; i < cpt.nbrDim(); ++i) {
144 const DiscreteVariable& pv = cpt.variable(i);
145 const auto [pbase, pslice] = _decode_(pv.name(), temporalSet);
147
149 pr.var = &pv;
150 pr.col = colByName[pbase];
151 pr.isAtemporal = pAtemporal;
153 nr.parents.push_back(std::move(pr));
154 }
155
156 _vars_[nr.col] = &v;
157 _inst_.add(v);
158 if (slice == lastSlice) _kernel_.push_back(_nodes_.size());
159 _nodes_.push_back(std::move(nr));
160 }
161 }
Size _nbVars_
number of base variable columns
std::vector< const DiscreteVariable * > _vars_
one representative variable per base column (same order as baseCols), pointing into template so it ou...
std::vector< std::string > _baseCols_
col index -> base name (canonical column numbering)
static std::pair< std::string, int > _decode_(const std::string &name, const std::unordered_set< std::string > &temporalSet)
decodes a template node name into (base, slice): "B[t]" with B a known temporal process -> (B,...
std::vector< Idx > _kernel_
indices, in nodes, of the slice-(k-1) nodes (the transition kernel, drives Phase 2); already in topol...
std::vector< NodeRef > _nodes_
all template nodes in topological order (drives Phase 1, the bootstrap)
Instantiation _inst_
a shared instantiation over all template variables, so we don't have to rebuild it for every draw
a template node, precompiled for fast sampling
a parent of a template node, precompiled for fast sampling

References _baseCols_, _decode_(), _inst_, _k_, _kernel_, _nbVars_, _nodes_, _template_, _vars_, gum::KTBN< GUM_SCALAR >::ATEMPORAL, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::NodeRef::col, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::ParentRef::col, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::NodeRef::cpt, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::ParentRef::isAtemporal, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::ParentRef::lag, gum::Variable::name(), gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::NodeRef::parents, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::NodeRef::slice, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::NodeRef::var, and gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::ParentRef::var.

Referenced by KTBNDatabaseGenerator().

Here is the call graph for this function:
Here is the caller graph for this function:

◆ _decode_()

template<GUM_Numeric GUM_SCALAR>
std::pair< std::string, int > gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_decode_ ( const std::string & name,
const std::unordered_set< std::string > & temporalSet )
staticprivate

decodes a template node name into (base, slice): "B[t]" with B a known temporal process -> (B, t); anything else (bare or bracket-named atemporal node) -> (name, ATEMPORAL).

Definition at line 89 of file KTBNDatabaseGenerator_tpl.h.

91 {
92 if (!name.empty() && name.back() == ']') {
93 const auto p = name.rfind('[');
94 if (p != std::string::npos && p + 1 < name.size() - 1) {
95 const std::string base = name.substr(0, p);
96 const std::string digits = name.substr(p + 1, name.size() - p - 2);
97 if (temporalSet.contains(base)
98 && std::all_of(digits.begin(), digits.end(), [](unsigned char c) {
99 return std::isdigit(c) != 0;
100 }))
101 return {base, std::stoi(digits)};
102 }
103 }
105 }

References gum::KTBN< GUM_SCALAR >::ATEMPORAL.

Referenced by _build_().

Here is the caller graph for this function:

◆ _drawSamples_()

template<GUM_Numeric GUM_SCALAR>
std::vector< double > gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_drawSamples_ ( Size nbSamples,
Size fixedLen,
const std::vector< Size > * perTraj,
std::string_view dirPath,
std::string_view csvBaseName,
VarOrderMode mode,
bool useLabels,
const std::string & csvSeparator )
private

the single worker behind both drawSamples() overloads. Trajectory i's horizon is read from perTraj (when non-null) else from fixedLen; samples each trajectory and writes it straight to its own CSV file.

The single worker behind both public drawSamples() overloads. Trajectory i's horizon T is read from perTraj when given, else from the shared fixedLen. Each trajectory is sampled into a flat row-major buffer (T × nbVars, in canonical column order) in two phases and written straight to its own CSV:

Phase 1 — initial k slices (0..k-1): no history yet, so every node of the k-slice template is drawn from scratch in topological order by inverse-CDF.

Phase 2 — transition (slices k..T-1): only the slice-(k-1) nodes (the kernel, already topological) are drawn, each parent read from the row at time (t - lag).

Definition at line 229 of file KTBNDatabaseGenerator_tpl.h.

236 {
237 // horizon of trajectory i: from perTraj when given, else the shared fixedLen
238 const auto lengthAt = [&](Idx i) { return perTraj ? (*perTraj)[i] : fixedLen; };
239
240 // validate everything up front, before any file is created
241 if (csvSeparator.find('\n') != std::string::npos)
242 GUM_ERROR(InvalidArgument, "csvSeparator must not contain end-line characters")
243 for (Idx i = 0; i < nbSamples; ++i)
244 if (lengthAt(i) < _k_)
245 GUM_ERROR(OperationNotAllowed, "nbTimeSlices=" << lengthAt(i) << " must be >= k=" << _k_)
246
247 // decide the column order once (shared by every trajectory file):
248 // colOrder[i] is the canonical column written at output position i.
250 switch (mode) {
254 default : GUM_ERROR(InvalidArgument, "unknown VarOrderMode")
255 }
256
259 std::filesystem::create_directories(dir); // ensure the destination exists
260
262 log2Ls.reserve(nbSamples);
263
264 const bool hasListener = onProgress.hasListener();
266 int progress = 0;
267 if (hasListener) {
268 timer.emplace();
269 GUM_EMIT2(onProgress, 0, 0.0);
270 }
271
272 // generate one trajectory at a time; each is sampled then written to its file
273 for (std::size_t i = 0; i < nbSamples; ++i) {
274 const Size T = lengthAt(i);
275
276 // flat trajectory buffer (row-major) in canonical column order: the value at
277 // time t, base column c is traj[t * _nbVars_ + c]. Sampling uses the cached
278 // integer columns (nr.col / pr.col) — no name lookup in the hot loop.
280 double log2L = 0;
281
282 // Phase 1: draw the first k time steps from scratch (no history available yet)
283 for (const NodeRef& nr: _nodes_) {
284 for (const ParentRef& pr: nr.parents) {
285 // an atemporal parent has the same value on every row, stored at row 0
286 const Size parentRow = pr.isAtemporal ? 0 : nr.slice - pr.lag;
287 _inst_.chgVal(*pr.var, traj[parentRow * _nbVars_ + pr.col]);
288 }
289 const Idx drawn = _drawVar_(*nr.var, *nr.cpt, log2L);
291 for (Size t = 0; t < T; ++t)
292 traj[t * _nbVars_ + nr.col] = drawn;
293 else traj[nr.slice * _nbVars_ + nr.col] = drawn;
294 }
295
296 // Phase 2: extend to T-1 by sliding the transition kernel forward one step at a time
297 for (Size t = _k_; t < T; ++t) {
298 for (const Idx kernelIdx: _kernel_) {
299 const NodeRef& nr = _nodes_[kernelIdx];
300 for (const ParentRef& pr: nr.parents) {
301 const Size parentRow = pr.isAtemporal ? 0 : t - pr.lag;
302 _inst_.chgVal(*pr.var, traj[parentRow * _nbVars_ + pr.col]);
303 }
304 traj[t * _nbVars_ + nr.col] = _drawVar_(*nr.var, *nr.cpt, log2L);
305 }
306 }
307
308 log2Ls.push_back(log2L);
309 const std::filesystem::path file = dir / (stem + std::to_string(i + 1) + ".csv");
311
312 if (hasListener) {
313 const int p = int((i * 100) / nbSamples);
314 if (p != progress) {
315 progress = p;
317 }
318 }
319 }
320
321 if (hasListener) {
323 ss << nbSamples << " trajectories generated in " << timer->step() << " s.";
324 GUM_EMIT1(onStop, ss.str());
325 }
326
327 return log2Ls;
328 }
Signaler< std::string_view > onStop
with a possible explanation for stopping
Signaler< Size, double > onProgress
Progression (percent) and time.
Idx _drawVar_(const DiscreteVariable &var, const Tensor< GUM_SCALAR > &cpt, double &log2likelihood)
inverse-CDF draw of var given the parents already set in inst; accumulates log2(P(drawn value)) into ...
void setVarOrderAntiTopological(std::vector< Idx > &colOrder) const
builds the reverse of setVarOrderTopological()
void setVarOrderTopological(std::vector< Idx > &colOrder) const
builds a topological column order (transition-kernel projection)
void _writeTrajectory_(std::string_view csvFileURL, const std::vector< Idx > &traj, Size nbTimeSlices, bool useLabels, const std::string &csvSeparator, const std::vector< Idx > &colOrder) const
writes one trajectory CSV (header + T rows) to csvFileURL. traj is the flat row-major buffer (T x nbV...
void setVarOrderRandomized(std::vector< Idx > &colOrder) const
builds a uniformly random column order

References _drawVar_(), _inst_, _k_, _kernel_, _nbVars_, _nodes_, _writeTrajectory_(), ANTI_TOPOLOGICAL, gum::KTBN< GUM_SCALAR >::ATEMPORAL, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::NodeRef::col, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::ParentRef::col, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::NodeRef::cpt, GUM_EMIT1, GUM_EMIT2, GUM_ERROR, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::ParentRef::isAtemporal, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::ParentRef::lag, gum::ProgressNotifier::onProgress, gum::ProgressNotifier::onStop, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::NodeRef::parents, RANDOM, setVarOrderAntiTopological(), setVarOrderRandomized(), setVarOrderTopological(), TOPOLOGICAL, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::NodeRef::var, and gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::ParentRef::var.

Referenced by drawSamples(), and drawSamples().

Here is the call graph for this function:
Here is the caller graph for this function:

◆ _drawVar_()

template<GUM_Numeric GUM_SCALAR>
Idx gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_drawVar_ ( const DiscreteVariable & var,
const Tensor< GUM_SCALAR > & cpt,
double & log2likelihood )
private

inverse-CDF draw of var given the parents already set in inst; accumulates log2(P(drawn value)) into log2likelihood

Definition at line 164 of file KTBNDatabaseGenerator_tpl.h.

166 {
167 const double threshold = gum::randomProba();
168 double cumulProb = 0.0;
169 for (_inst_.setFirstVar(var); !_inst_.end(); _inst_.incVar(var)) {
170 cumulProb += cpt[_inst_];
171 if (cumulProb > threshold) break;
172 }
173 if (_inst_.end()) _inst_.setLastVar(var);
175 return _inst_.val(var);
176 }
GUM_SHARED_PUBLIC double randomProba()
Returns a random double between 0 and 1 included (i.e.

References _inst_, and gum::randomProba().

Referenced by _drawSamples_().

Here is the call graph for this function:
Here is the caller graph for this function:

◆ _label_()

template<GUM_Numeric GUM_SCALAR>
std::string gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_label_ ( Idx col,
Idx idx ) const
private

renders the label of modality idx of base column col (taking the discretized-label mode into account)

Definition at line 346 of file KTBNDatabaseGenerator_tpl.h.

346 {
347 const DiscreteVariable& v = *_vars_[col];
348 if (v.varType() == VarType::DISCRETIZED) {
349 switch (_discretizedLabelMode_) {
350 case DiscretizedLabelMode::MEDIAN : return std::to_string(v.numerical(idx));
352 return std::to_string(static_cast< const IDiscretizedVariable& >(v).draw(idx));
353 case DiscretizedLabelMode::INTERVAL : return v.label(idx);
354 default : GUM_ERROR(FatalError, "unknown DiscretizedLabelMode")
355 }
356 }
357 return v.label(idx);
358 }
DiscretizedLabelMode _discretizedLabelMode_
rendering of discretized variables when labels are requested

References _discretizedLabelMode_, _vars_, gum::DISCRETIZED, GUM_ERROR, INTERVAL, gum::DiscreteVariable::label(), MEDIAN, gum::DiscreteVariable::numerical(), RANDOM, and gum::DiscreteVariable::varType().

Referenced by _writeTrajectory_().

Here is the call graph for this function:
Here is the caller graph for this function:

◆ _writeTrajectory_()

template<GUM_Numeric GUM_SCALAR>
void gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_writeTrajectory_ ( std::string_view csvFileURL,
const std::vector< Idx > & traj,
Size nbTimeSlices,
bool useLabels,
const std::string & csvSeparator,
const std::vector< Idx > & colOrder ) const
private

writes one trajectory CSV (header + T rows) to csvFileURL. traj is the flat row-major buffer (T x nbVars) in canonical column order; colOrder gives the output column order.

Definition at line 361 of file KTBNDatabaseGenerator_tpl.h.

367 {
369 if (!os) GUM_ERROR(IOError, "could not open '" << csvFileURL << "' for writing")
370
371 // header: one column per base variable, in the chosen output order
374 if (!firstCol) os << csvSeparator;
375 os << _baseCols_[col];
376 firstCol = false;
377 }
378 os << "\n";
379
380 for (Size t = 0; t < nbTimeSlices; ++t) {
381 const Size base = t * _nbVars_;
382 firstCol = true;
383 for (const Idx col: colOrder) {
384 if (!firstCol) os << csvSeparator;
386 firstCol = false;
387 }
388 os << "\n";
389 }
390 }
std::string _label_(Idx col, Idx idx) const
renders the label of modality idx of base column col (taking the discretized-label mode into account)

References _baseCols_, _label_(), _nbVars_, and GUM_ERROR.

Referenced by _drawSamples_().

Here is the call graph for this function:
Here is the caller graph for this function:

◆ drawSamples() [1/2]

template<GUM_Numeric GUM_SCALAR>
std::vector< double > gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::drawSamples ( const std::vector< Size > & nbTimeSlices,
std::string_view dirPath,
std::string_view csvBaseName,
VarOrderMode mode = VarOrderMode::RANDOM,
bool useLabels = true,
std::string csvSeparator = "," )

Like drawSamples(), but every trajectory may have its own horizon.

Parameters
nbTimeSlicesOne horizon per trajectory; its size is the number of trajectories. Every entry must be \(\geq k\).
dirPathDirectory to write the CSV files into.
csvBaseNameStem for each file name (index and .csv appended).
modeThe column order of the base variables.
useLabelsRender values as variable labels (else modality index).
csvSeparatorColumn separator (must not contain a newline).
Returns
The log2-likelihood of each generated trajectory.
Exceptions
OperationNotAllowedif some entry of nbTimeSlices is smaller than \(k\).

Definition at line 200 of file KTBNDatabaseGenerator_tpl.h.

205 {
206 // per-trajectory horizons: nbTimeSlices[i] is trajectory i's length
207 return _drawSamples_(nbTimeSlices.size(),
208 0,
210 dirPath,
212 mode,
213 useLabels,
215 }
std::vector< double > _drawSamples_(Size nbSamples, Size fixedLen, const std::vector< Size > *perTraj, std::string_view dirPath, std::string_view csvBaseName, VarOrderMode mode, bool useLabels, const std::string &csvSeparator)
the single worker behind both drawSamples() overloads. Trajectory i's horizon is read from perTraj (w...

References _drawSamples_().

Here is the call graph for this function:

◆ drawSamples() [2/2]

template<GUM_Numeric GUM_SCALAR>
std::vector< double > gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::drawSamples ( Size nbSamples,
Size nbTimeSlices,
std::string_view dirPath,
std::string_view csvBaseName,
VarOrderMode mode = VarOrderMode::RANDOM,
bool useLabels = true,
std::string csvSeparator = "," )

Generates nbSamples independent trajectories, writing one CSV file per trajectory into dirPath.

File names are csvBaseName followed by the 1-based trajectory index and ".csv" (e.g. "traj1.csv", "traj2.csv", …). Each file has one column per base variable in the order chosen by mode, and nbTimeSlices data rows.

Parameters
nbSamplesThe number of trajectories to generate.
nbTimeSlicesThe horizon \(T\) shared by every trajectory. Must be \(\geq k\).
dirPathDirectory to write the CSV files into.
csvBaseNameStem for each file name (index and .csv appended).
modeThe column order of the base variables.
useLabelsRender values as variable labels (else modality index).
csvSeparatorColumn separator (must not contain a newline).
Returns
The log2-likelihood of each generated trajectory (size nbSamples).
Exceptions
OperationNotAllowedif nbTimeSlices is smaller than \(k\).

Definition at line 180 of file KTBNDatabaseGenerator_tpl.h.

186 {
187 // fixed horizon: every trajectory shares nbTimeSlices (perTraj == nullptr)
190 nullptr,
191 dirPath,
193 mode,
194 useLabels,
196 }

References _drawSamples_().

Here is the call graph for this function:

◆ nbVars()

template<GUM_Numeric GUM_SCALAR>
INLINE Size gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::nbVars ( ) const

returns the number of base variable columns

Definition at line 72 of file KTBNDatabaseGenerator_tpl.h.

72 {
73 return _nbVars_;
74 }

References _nbVars_.

◆ operator=() [1/2]

template<GUM_Numeric GUM_SCALAR>
KTBNDatabaseGenerator & gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::operator= ( const KTBNDatabaseGenerator< GUM_SCALAR > & )
privatedelete

References KTBNDatabaseGenerator().

Here is the call graph for this function:

◆ operator=() [2/2]

template<GUM_Numeric GUM_SCALAR>
KTBNDatabaseGenerator & gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::operator= ( KTBNDatabaseGenerator< GUM_SCALAR > && )
privatedelete

References KTBNDatabaseGenerator().

Here is the call graph for this function:

◆ setDiscretizedLabelModeInterval()

template<GUM_Numeric GUM_SCALAR>
void gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::setDiscretizedLabelModeInterval ( )

set discretized-label rendering to the interval label "[min,max["

Definition at line 341 of file KTBNDatabaseGenerator_tpl.h.

References _discretizedLabelMode_, and INTERVAL.

◆ setDiscretizedLabelModeMedian()

template<GUM_Numeric GUM_SCALAR>
void gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::setDiscretizedLabelModeMedian ( )

set discretized-label rendering to the (deterministic) interval median

Definition at line 336 of file KTBNDatabaseGenerator_tpl.h.

References _discretizedLabelMode_, and MEDIAN.

◆ setDiscretizedLabelModeRandom()

template<GUM_Numeric GUM_SCALAR>
void gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::setDiscretizedLabelModeRandom ( )

set discretized-label rendering to a uniform random draw in the interval (this is the default; each labelled export then differs)

Definition at line 331 of file KTBNDatabaseGenerator_tpl.h.

References _discretizedLabelMode_, and RANDOM.

◆ setVarOrderAntiTopological()

template<GUM_Numeric GUM_SCALAR>
void gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::setVarOrderAntiTopological ( std::vector< Idx > & colOrder) const
private

builds the reverse of setVarOrderTopological()

Definition at line 460 of file KTBNDatabaseGenerator_tpl.h.

461 {
463 std::reverse(colOrder.begin(), colOrder.end());
464 }

References setVarOrderTopological().

Referenced by _drawSamples_().

Here is the call graph for this function:
Here is the caller graph for this function:

◆ setVarOrderRandomized()

template<GUM_Numeric GUM_SCALAR>
void gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::setVarOrderRandomized ( std::vector< Idx > & colOrder) const
private

builds a uniformly random column order

Definition at line 393 of file KTBNDatabaseGenerator_tpl.h.

394 {
395 colOrder.resize(_nbVars_);
396 std::iota(colOrder.begin(), colOrder.end(), 0);
398 }

References _nbVars_, and gum::randomGenerator().

Referenced by _drawSamples_().

Here is the call graph for this function:
Here is the caller graph for this function:

◆ setVarOrderTopological()

template<GUM_Numeric GUM_SCALAR>
void gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::setVarOrderTopological ( std::vector< Idx > & colOrder) const
private

builds a topological column order (transition-kernel projection)

Definition at line 401 of file KTBNDatabaseGenerator_tpl.h.

402 {
403 // The column order is a topological sort of the "contemporaneous" DAG over
404 // base columns. Only two kinds of arc constrain the column order:
405 // * atemporal -> atemporal — orders the atemporal block
406 // * same-slice kernel arcs — lag-0 arcs at slice k-1; order the temporal block
407 // Lagged arcs are resolved at sampling time, not here; atemporal -> temporal
408 // arcs need no edge since every atemporal column already precedes every
409 // temporal one (the blocks are emitted in that order below).
413
414 auto addEdge = [&](Idx parent, Idx child) {
415 children[parent].push_back(child);
416 ++indegree[child];
417 };
418
419 // (1) build the contemporaneous DAG, one arc set per block.
420 const int lastSlice = int(_k_) - 1;
421 for (const NodeRef& nf: _nodes_) {
422 if (nf.slice == KTBN< GUM_SCALAR >::ATEMPORAL) {
423 isAtemporal[nf.col] = true;
424 for (const ParentRef& pr: nf.parents) {
425 if (!pr.isAtemporal)
427 "temporal variable '"
428 << _baseCols_[pr.col] << "' is a parent of atemporal variable '"
429 << _baseCols_[nf.col] << "' — violates the k-DBN invariant")
430 addEdge(pr.col, nf.col);
431 }
432 } else if (nf.slice == lastSlice) {
433 for (const ParentRef& pr: nf.parents)
434 if (!pr.isAtemporal && pr.lag == 0) addEdge(pr.col, nf.col);
435 }
436 }
437
438 // (2) Kahn's sort of one block, seeded in ascending col order so the result
439 // is deterministic (the seed is canonical and edges follow _nodes_).
440 auto topoSortBlock = [&](bool atempBlock) {
442 for (Idx c = 0; c < _nbVars_; ++c)
443 if (isAtemporal[c] == atempBlock && indegree[c] == 0) q.push(c);
444 while (!q.empty()) {
445 const Idx cur = q.front();
446 q.pop();
447 colOrder.push_back(cur);
448 for (const Idx ch: children[cur])
449 if (--indegree[ch] == 0) q.push(ch);
450 }
451 };
452
453 // (3) atemporal block first, then the temporal kernel order.
454 colOrder.clear();
455 topoSortBlock(true);
456 topoSortBlock(false);
457 }

References _baseCols_, _k_, _nbVars_, _nodes_, gum::KTBN< GUM_SCALAR >::ATEMPORAL, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::ParentRef::col, GUM_ERROR, gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::ParentRef::isAtemporal, and gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::ParentRef::lag.

Referenced by _drawSamples_(), and setVarOrderAntiTopological().

Here is the caller graph for this function:

Member Data Documentation

◆ _baseCols_

template<GUM_Numeric GUM_SCALAR>
std::vector< std::string > gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_baseCols_
private

col index -> base name (canonical column numbering)

Definition at line 227 of file KTBNDatabaseGenerator.h.

Referenced by _build_(), _writeTrajectory_(), and setVarOrderTopological().

◆ _discretizedLabelMode_

template<GUM_Numeric GUM_SCALAR>
DiscretizedLabelMode gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_discretizedLabelMode_ = DiscretizedLabelMode::RANDOM
private

rendering of discretized variables when labels are requested

Definition at line 245 of file KTBNDatabaseGenerator.h.

Referenced by _label_(), setDiscretizedLabelModeInterval(), setDiscretizedLabelModeMedian(), and setDiscretizedLabelModeRandom().

◆ _inst_

template<GUM_Numeric GUM_SCALAR>
Instantiation gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_inst_
private

a shared instantiation over all template variables, so we don't have to rebuild it for every draw

Definition at line 242 of file KTBNDatabaseGenerator.h.

Referenced by _build_(), _drawSamples_(), and _drawVar_().

◆ _k_

template<GUM_Numeric GUM_SCALAR>
Size gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_k_
private

the order \(k\) of the k-DBN

Definition at line 221 of file KTBNDatabaseGenerator.h.

Referenced by KTBNDatabaseGenerator(), _build_(), _drawSamples_(), and setVarOrderTopological().

◆ _kernel_

template<GUM_Numeric GUM_SCALAR>
std::vector< Idx > gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_kernel_
private

indices, in nodes, of the slice-(k-1) nodes (the transition kernel, drives Phase 2); already in topological order

Definition at line 238 of file KTBNDatabaseGenerator.h.

Referenced by _build_(), and _drawSamples_().

◆ _nbVars_

template<GUM_Numeric GUM_SCALAR>
Size gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_nbVars_
private

number of base variable columns

Definition at line 224 of file KTBNDatabaseGenerator.h.

Referenced by _build_(), _drawSamples_(), _writeTrajectory_(), nbVars(), setVarOrderRandomized(), and setVarOrderTopological().

◆ _nodes_

template<GUM_Numeric GUM_SCALAR>
std::vector< NodeRef > gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_nodes_
private

all template nodes in topological order (drives Phase 1, the bootstrap)

Definition at line 234 of file KTBNDatabaseGenerator.h.

Referenced by _build_(), _drawSamples_(), and setVarOrderTopological().

◆ _template_

template<GUM_Numeric GUM_SCALAR>
BayesNet< GUM_SCALAR > gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_template_
private

the \(k\)-slice template (a small copy, independent of the horizon)

Definition at line 218 of file KTBNDatabaseGenerator.h.

Referenced by KTBNDatabaseGenerator(), and _build_().

◆ _vars_

template<GUM_Numeric GUM_SCALAR>
std::vector< const DiscreteVariable* > gum::learning::KTBNDatabaseGenerator< GUM_SCALAR >::_vars_
private

one representative variable per base column (same order as baseCols), pointing into template so it outlives the source k-DBN. Label rendering only.

Definition at line 231 of file KTBNDatabaseGenerator.h.

Referenced by _build_(), and _label_().

◆ onProgress

Signaler< Size, double > gum::ProgressNotifier::onProgress
inherited

◆ onStop

Signaler< std::string_view > gum::ProgressNotifier::onStop
inherited

The documentation for this class was generated from the following files: