This repository contains a deterministic, config-driven generator for synthetic telecom-like Call Detail Record (CDR) event streams. It is designed as a reusable research artifact for experiments involving fraud injection, abrupt drift, and sequential stream-style consumption.
The repository intentionally contains only the generator and its reference configuration. Experiment pipelines, detector benchmarking, plots, and paper assets belong in a separate downstream repository.
- Deterministic output under a fixed random seed
- Time-ordered telecom-like CDR events
- Configurable fraud prevalence before and after drift
- Abrupt drift injection at a known timestamp
- Caller-history features such as recent activity and international ratio
- CSV export by default, with optional Parquet export
- Batch iterator support for stream-style consumption
.
├── configs/
│ └── default.yaml
├── data/
│ └── demo_stream.csv
├── src/
│ └── telecom_stream/
├── generate_stream.py
├── requirements.txt
└── CITATION.cff
pip install -r requirements.txtIf you want Parquet output, install pyarrow separately.
Generate the reference demo dataset:
python generate_stream.py --config configs/default.yaml --output data/demo_stream.csvGenerate Parquet instead:
python generate_stream.py --config configs/default.yaml --output data/demo_stream.parquetIf --output is omitted, the CLI writes to data/generated_stream.<default_format> using the default format from the config file.
Generated events include:
event_idevent_timecaller_idcallee_idcall_duration_seccall_typedestination_typeregionroaming_flagsim_age_dayscalls_last_24hunique_callees_last_24hinternational_ratio_7dfraud_labeldrift_regimehour_of_dayday_of_weekis_weekendnetwork_type
The default configuration injects an abrupt drift at a configurable timestamp.
- Before drift: fraud is shorter, burstier, and more international.
- After drift: fraud becomes less obviously international and closer to normal traffic while remaining behaviorally distinct.
The active regime is stored in drift_regime as baseline or post_drift.
The YAML config controls:
- random seed
- subscriber and event counts
- time range
- fraud ratios before and after drift
- drift timestamp
- behavior profiles for normal, pre-drift fraud, and post-drift fraud
- region choices
- export defaults and batch size
See configs/default.yaml for the reference configuration.
Repository metadata is provided in CITATION.cff. After creating a GitHub release and archiving it in Zenodo, update the citation file with the release DOI for paper-ready citation.
Add a license file before making the repository public for reuse.