Automated data cleaning · built in-house

Every respondent scored, every exclusion explained.

ResClean is the respondent quality engine behind every MIU dataset. It reads your raw export, auto-detects the questionnaire, runs 24 fraud and quality checks and returns one auditable Data Quality Score per interview — with the evidence that produced it. It runs inside your browser, so the file never reaches a server.

24fraud & quality checks per respondent
6+1sub-scores rolled into one Data Quality Score
0files uploaded — processing is local
The score system

Six dimensions. One number you can defend.

Each sub-score runs 0–100 where 100 is clean. You set the weights and the Keep / Review / Remove cut-offs, and both are written into every export so the run can be reproduced on the next wave.

CHSCoherence Score

Flat grids, near-flat grids, zig-zag, staircase and repeating-cycle patterning, extreme and midpoint response styling, non-differentiation against the sample and reverse-keyed contradictions.

TSTime Score

Interview length against the sample median, hyper-speeding, seconds per question, seconds per open end, Qualtrics-style per-page timers and implausibly long idle sessions.

OESOpen-Ended Score

Keyboard mash and non-language text, AI-written prose, low-effort filler, off-topic answers, verbatims duplicated across respondents and paste detection from typing speed.

BASBehavioural Analytics Score

Batch entry with machine-regular arrival gaps, duplicate identifiers, near-identical answer vectors, repeat IPs, /24 subnet clusters, click-farm device signatures and geo mismatches.

CSCompletion Score

How much of the questionnaire the respondent actually answered, measured across the questions you selected rather than the whole file.

ACAttention Check Score

Pass rate on the trap and instructed-response items you designate. Codes and answer labels both work, and several acceptable answers can be allowed.

DQSData Quality Score

The weighted roll-up. Under your removal cut-off the record is dropped; between the cut-offs it is retained but flagged for a human look; a single critical finding always earns review, whatever the mean says.

What it catches

Twenty-four checks across six families.

No single rule proves fraud. ResClean combines behavioural, statistical, textual and metadata signals, then reports which ones fired and which could not run because a column was not available.

Pattern

  • Straightlining
  • Near-straightlining
  • Zig-zag / staircase / cycles
  • Extreme response style
  • Midpoint response style
  • Reverse-item contradictions

Timing

  • Speeders
  • Hyper-speeders
  • Per-question pace floor
  • Per-page speeding
  • Pasted verbatims
  • Implausibly long sessions

Open text

  • Gibberish / keyboard mash
  • Likely AI-written prose
  • Low-effort filler
  • Blank verbatims
  • Off-topic answers
  • Duplicated verbatims

Fraud & duplicates

  • Duplicate identifiers
  • Near-identical answer vectors
  • Batch entrants
  • Repeat IP addresses
  • Subnet clusters
  • Click-farm signatures
  • Geo mismatch
  • Failed attention checks
  • High item non-response
Why ResClean

Where it goes further than a generic quality tool.

Most platforms score the answers they can parse and ask you to trust the verdict. ResClean is built for research operations: it reads production export formats, understands survey structure, and hands you the audit trail.

Comparison of manual review, a generic quality platform and ResClean
CapabilityManual review in Excel/SPSSGeneric quality platformResClean by Miures
Native SPSS .sav inputNeeds SPSS licenceUsually CSV / XLSX onlyRead directly — labels, measure levels, long strings, all compression modes
Question-type detectionManual mappingColumn-level guessesGrids, multi-punch, ranking, numeric and open ends grouped by stem, with confidence and reason
Multi-header exportsManual clean-upOften breaksQualtrics three-row and Decipher two-row headers detected automatically
Codes and labels togetherTwo files, manual lookupCodes onlyPair a labels export with a codes export and read real answer text everywhere
AI-written verbatim detectionAnalyst judgementScore with no reasonsScore plus the specific triggers — phrasing, structure, register, sentence uniformity, typing speed
Fraud clusteringPivot tables, if attemptedDuplicates and IPsBatch-entry gap regularity, /24 subnets, device + duration signatures, near-identical answer vectors
Audit trailWhatever was written downScore exportFlags, sub-scores, reason text, thresholds used and a printable methodology report
Tracker reuseRe-specified each waveProject-level settingsSaved cleaning profile — thresholds, question map, traps and reverse-keyed items — reapplied in one click
Data residencyYour machineVendor cloudYour browser. Nothing is uploaded, stored or logged
Who decidesYouAutomated thresholdYou. A critical flag triggers review, never a silent delete, and any disposition can be overridden

Comparison reflects MIU's own capability and the general shape of manual and platform-based alternatives. Individual vendors differ — ask us what a specific tool does and does not cover before switching.

Interactive visual demo

From raw export to an auditable clean dataset.

Four stages, each producing a file you can inspect. Nothing happens off-screen.

Ready for a project-specific estimate?

Open the enquiry form with this service selected, review the prefilled details and click Send enquiry for review.

FileRowsFormattracker_w3.sav1,240SPSSqualtrics_raw.csv8603 header rows
Stage 1 of 4

Read the export as it actually comes out

Value and label structure, multi-row headers and SPSS dictionaries are parsed locally — no reformatting first.

  • CSV, TSV, XLSX, .sav, .zsav
  • Qualtrics and Decipher header rows
  • Codes paired with a labels export
1 / 4
Service tiers

Buy the engine, or the engine plus our analysts.

Quoted per completed interview with a project minimum. Add open-ended screening per record, and take a discount from wave 2 of a tracker when the saved profile is reused.

From $0.30 / complete

Automated screening

The ResClean run, the scored dataset and the quality report. You make the exclusion calls.

  • Six sub-scores and a DQS per interview
  • Flag file with reason codes
  • Printable methodology report
  • 1 business-day planning TAT
From $0.55 / complete

Managed cleaning

Everything in screening, plus an analyst reviewing every flagged record and cluster before delivery.

  • Human review of all consequential exclusions
  • Client-agreed thresholds and decision log
  • Retained and removed datasets delivered separately
  • 2 business-day planning TAT
From $0.85 / complete

Forensic cleaning

For disputed data, suspect suppliers or trackers where the removal rate has to be defended.

  • Supplier / source quality scorecards
  • Wave-over-wave incidence comparison
  • Cluster investigation and recontact review where evidence allows
  • 3 business-day planning TAT
What you receive

Files a client or committee can audit.

  • Cleaned workbook: original columns plus flags, sub-scores, DQS and reason text.
  • Retained and removed records as separate sheets, never silently dropped.
  • Flag matrix — one row per respondent, one column per check.
  • Question map showing what was scored and what was excluded from scoring.
  • Quality report with disposition split, score distribution, flag incidence and cluster table.
  • The exact thresholds and weights used, so the wave can be reproduced.
  • Reusable cleaning profile as JSON.
Honest limits

What we will not claim.

  • No tool can guarantee zero fraud. Fraud evolves and legitimate respondents sometimes look unusual.
  • AI-verbatim detection is probabilistic. It reports its reasons and is meant to prompt a read, not to prove authorship.
  • Checks that need an unmapped column are skipped — and the report says so rather than implying a clean bill of health.
  • Geo mismatch needs both a stated and an IP-derived country.
  • Final confidence still depends on recruitment source, respondent verification and questionnaire design.
Questions

The things buyers ask first.

Does my file get uploaded?

No. ResClean parses and scores it inside your own browser tab. Nothing is uploaded, stored or logged, so respondent PII never leaves your machine — and there is no server-imposed file-size limit.

Which formats can it read?

CSV, TSV, Excel .xlsx and SPSS .sav / .zsav, including value labels, measure levels, long variable names and very long strings. Qualtrics three-row and Decipher two-row header exports are detected automatically.

Can it really spot AI answers?

It scores AI-likeness from assistant phrasing, markdown structure, stacked essay connectives, metronomic sentence lengths, impersonal register, length against the sample norm and typing speed. The triggers are reported with the score. Read the verbatim before acting.

Will it delete data automatically?

Never silently. Each record gets Keep, Review or Remove from thresholds you control, every decision carries its reason, and you can override any of them before export. Removed records are delivered, not discarded.

Can you clean a study you did not run?

Yes — cleaning is a standalone line item. Send the raw export, the questionnaire and any known concerns; it works regardless of who programmed or fielded the study.

What about tracker waves?

Save the thresholds, question map, trap items and reverse-keyed items as a profile and reapply it to the next wave in one click — so wave-on-wave removal rates are comparable rather than re-specified.

Start with your own data

Run ResClean on a real file before you talk to anyone.

Open the engine, drop in an export and look at the flags. If you want our analysts on it, send the questionnaire, the raw file and the quality concerns you already suspect.