Flagship Build — Research-grade Survey Data Analysis

Survey & ODK Paradata Platform

A research-grade data-analysis platform that turns raw SurveyCTO and ODK paradata — millions of per-event timing rows — into 40 engineered features per interview, scores each one for fraud and quality anomalies with robust IQR statistics, and surfaces the whole thing in an analytics dashboard evaluated against a ground-truth label set.

PythonFastAPIPandasNumPyMatplotlibReactViteZustandPytest
Engineered features40
DetectionRobust IQR
F1 vs ground truth76.1%
Validationpytest E2E
Survey paradata analytics dashboard with anomaly KPIs, recall, and F1 score

Role focus

Data engineering, feature extraction, statistical anomaly detection, and analytics dashboard

Pipeline testsDetection recall

Project narrative

This is the strongest data-analysis story in the archive because it treats survey quality as a measurable, testable problem rather than a spreadsheet chore. Field survey data is easy to fake — an enumerator speed-runs a 30-minute interview in two minutes, or straightlines every answer — and the tell is in the timing. The platform reads the audit log every survey app already records (every question appearance, button press, and back-navigation), mathematically compresses those millions of rows into one predictive row per interview, and flags the fakes. It is deliberately honest: it reports precision, recall, and F1 against a labelled adversarial set instead of a single flattering number, and it is backed by the most thorough automated-testing suite in the archive.

Why it matters

Research-grade data engineering: feature extraction, applied statistics (entropy, skewness, kurtosis, IQR), and honest evaluation against ground truth

A complete story from raw paradata → engineered features → anomaly scoring → an analytics dashboard

The most disciplined automated-testing practice in the archive, backed by an academic report and reproducible metrics

Preview Gallery

Visuals that make the project story land faster.

Interface views and system visuals that help the product story read quickly. Tap any screenshot to view it full size.

Architecture map

Pipeline

A four-layer Python validation engine: structural ingestion (CIR parity, strict CSV schema), an integrity layer that traps logically impossible timings, a behavioral layer that extracts 40 features (entropy, speed trajectories, back-navigation, straightlining), and a classification layer that maps weighted scores to a three-tier severity using robust IQR bounds.

Service and data

A FastAPI service exposes datasets and results, runs the pipeline on demand, and can generate synthetic adversarial audit logs; a strict CSV contract is enforced on ingest, and a ground-truth label set drives honest evaluation.

Analytics dashboard

A React + Vite dashboard with an overview of KPIs and anomaly breakdowns, a flagged-anomaly table, per-feature distributions and statistics, an audit report with the pipeline's output plots, and a raw feature-matrix explorer.

Engineering Decisions

This is where the software engineering depth shows up.

Beyond features, this section highlights the structural choices that shape scalability, reliability, and product clarity.

Chose robust IQR bounds over naive Z-scores because survey timing data is heavily right-skewed and non-normal, so mean-and-standard-deviation thresholds would misfire on real distributions.
Separated the engine into four layers — ingestion, integrity, behavioral, classification — so mathematical integrity is enforced before any feature is trusted, and a bad row fails early rather than silently corrupting a score.
Evaluated against a labelled adversarial ground-truth set and reported precision, recall, and F1 honestly, rather than shipping a single flattering accuracy number.

Validation snapshot

Pipeline tests

An end-to-end pytest suite with statistical threshold assertions — the strongest automated-testing story in the archive.

Detection recall

Precision is 100% (it does not cry wolf), but recall is 61.4% against ground truth — it currently catches 27 of 44 known-bad interviews and misses slower-burn anomalies, which the evaluation reports openly rather than hiding.