Output & Determinism
Output Formats & Determinism
LotusGen enforces deterministic generation and high-throughput data exports.
Reproducibility via Master Seeds
Every project in LotusGen defines an integer seed. All pseudorandom processes—including UUIDv7 identifiers, weighted sampling, regex generators, and normal distribution sampling—derive from this seed.
Re-running a generation job with identical configuration and seed produces byte-for-byte identical output datasets.
Supported File Formats
1. CSV (Comma-Separated Values)
- Standard UTF-8 encoded text format.
- Streamed in row batches to prevent excessive buffer allocation.
- Compatible with all standard spreadsheet and ETL tooling.
2. Apache Parquet
- Open-source columnar storage format optimized for analytics and data warehousing.
- Includes embedded schema definitions and Snappy compression.
- Up to 80-90% smaller disk footprint compared to raw CSV.
- Ideal for ingestion into DuckDB, ClickHouse, Apache Spark, and Pandas/Polars dataframes.