WiFi RSSI mobility classification: infer where a passenger is (inside the bus, at the stop) and which transition they are making, from signal-strength sequences captured by a fixed access point.
Part of the inMotion project (Integrated AI Platform for Urban Mobility and Transportation), a copromotion between Wavecom and Instituto de Telecomunicações.
Route labels — each row is one 10-second route, 10 RSSI readings at 1 Hz:
| label | meaning |
|---|---|
AA |
stayed inside the bus |
BB |
stayed at the stop |
AB |
alighted (inside → outside) |
BA |
boarded (outside → inside) |
Each row also carries mac (device identifier) and noise (whether other
devices were transmitting concurrently). See docs/eucnc/project_description.md
for the collection protocol.
uv sync --all-packages # root project plus the demo workspace member
cp env.example .env # credentials (never committed)
set -a && . ./.env && set +aCheck the environment and the artifact store:
uv run inmotion doctorThis is the part that used to require reading the training scripts.
uv run inmotion model list # every known model, and whether it is on disk
uv run inmotion model info ts_jepa # architecture, parameter count, provenance
uv run inmotion model verify # strict-load all of them
uv run inmotion model predict ts_jepa --input new_routes.csvmodel predict needs to know how the inputs were scaled, because the original
training code never saved its StandardScaler. Pass the CSV the model was
trained on as --reference, or rely on the default for its feature set.
Predictions built on reconstructed statistics say so.
A checkpoint that predates the sidecar convention can be adopted:
uv run inmotion model adopt path/to/checkpoint.pt --name my_model
uv run inmotion model inventory 20-aug # scan a whole directoryAdoption rebuilds the architecture from the weights and refuses to guess when several architectures fit (TS-JEPA and CF-JEPA at equal width are genuinely indistinguishable). It reports the ambiguity instead.
Note:
*_pretrain_best.ptfiles are pre-training bundles, not classifiers. Seedocs/reports/ARTIFACTS.md.
Datasets are committed to git, so their history is the version history.
data/DATASETS.json adds the parts git cannot express: checksums, row and class
counts, lineage, and which versions are safe to train on.
uv run inmotion datasets list
uv run inmotion datasets show ds-augmented-v3
uv run inmotion datasets verify # detect a file changed under git's feetCurrent versions, newest lineage last:
| id | file | rows | notes |
|---|---|---|---|
ds-pure |
dataset_only_pure.csv |
160 | early collection — labels reported incorrect |
ds-noise |
dataset_only_noise.csv |
1,196 | early collection — labels reported incorrect |
ds-augmented-v1 |
dataset_augmented.csv |
10,113 | augmentation config 1 — labels reported incorrect |
ds-icaisf |
dataset.icaisf.csv |
2,251 | the ISAC paper's variant; labels unverified |
ds-base-v2 |
dataset.csv |
3,511 | current base, 13 devices |
ds-augmented-v2 |
dataset_augmented2.csv |
20,275 | augmentation config 2 (superseded) |
ds-augmented-v3 |
dataset_augmented3.csv |
35,839 | canonical training set |
ds-augmented-v3 is what the JEPA and SIGReg checkpoints were trained on, so it
is the reference for reproducing their preprocessing. Full detail, including
lineage and checksums, is in docs/reference/datasets.md.
Building datasets from raw captures:
uv run inmotion data export --input wavecom_files/routeAtoB.txt --output data/AB.csv \
--group G1:<mac1>,<mac2> --group-label G1:AB
uv run inmotion data merge --input data --output dataset.csv
uv run inmotion data clean -- --data dataset.csv
uv run inmotion data augment -- ...Pipelines are invoked through one entry point; arguments pass through unchanged.
uv run inmotion run list
uv run inmotion train exotic --model ts_jepa --seed 42
uv run inmotion train dl -- --seed 42
uv run inmotion evaluate mega-ensemble --data dataset_augmented3.csvCluster submission scripts are in scripts/slurm/, single-machine runs in
scripts/local/; see scripts/README.md.
uv run pytest tests/ # root suite
cd demo && uv run --project . pytest # demo is a separate workspace member
uv run ruff check src dl ml_classification tests
uv run ruff format --check src dl ml_classification tests
uv run python scripts/docs/generate_reference.py --checkscripts/check.sh runs all of them at once:
./scripts/check.sh # lint, format, docs, tests, rust checks
./scripts/check.sh --full # also the demo suite and the Rust parity testThere is deliberately no CI: the checks run locally, on demand, so they cost nothing.
Model tests verify that every documented checkpoint loads with strict=True and
matches its recorded parameter count, and that each model's output still matches
a frozen fingerprint. They skip with a reason when the (gitignored) checkpoint
store is absent.
src/inmotion/ the package
cli.py the `inmotion` entry point
models/ registry, factories, sidecars, adoption
data/ tier paths, feature pipeline, export/merge/clean/augment
pipelines/ training and evaluation pipelines + their registry
evaluation/ checkpoint evaluation, SHAP, feature importance
analysis/ viz/ multi-seed analysis, figure regeneration
data/
raw/ interim/ processed/ the three data tiers
DATASETS.json checksums, lineage, label provenance
docs/ papers (frozen), reference (generated), guides, audit
scripts/ slurm/ local/ archive/ oneoff/ docs/
tests/ contract, adoption, CLI, docs, secrets guard
demo/ FastAPI demo (separate workspace member, own image)
backup/ 20-aug/ checkpoint stores (not tracked)
plots/ results/ generated output (not tracked)
For a review, these are the entry points: docs/audit/AUDIT_CHECKLIST.md
lists every claim and the command that verifies it, and
docs/audit/MODEL_CARDS.md states each model's
purpose and limitations. Known discrepancies — including three checkpoints whose
parameter counts do not match the published table — are documented in
docs/reproducibility.md rather than omitted.
docs/README.md is the documentation index. Start there.
| document | contents |
|---|---|
docs/architecture.md |
how data flows, the model families, the registry, sidecars |
docs/reproducibility.md |
environment, seeds, split, and what is not reproducible |
docs/guides/running-models.md |
list, inspect, verify, predict, adopt |
docs/guides/training.md |
training each family, on a workstation or a cluster |
docs/reference/ |
generated: CLI, models, modules, datasets |
docs/audit/AUDIT_CHECKLIST.md |
what a reviewer can verify independently |
docs/audit/MODEL_CARDS.md |
per-family model cards and limitations |
docs/reports/ |
artifact rules, migration notes, refactor history |
Dataset DOI: 10.21227/55nm-0r91
