BLINDAUDIOTESTLISTEN. COMPARE. DECIDE.

.batest Specification

This page renders a copy of the .batest file format specification, pinned to release v2.6. View the full specification and changelog on GitHub for the authoritative, versioned source.

Blind Audio Test File Format Specification (.batest)

Version: 2.6 Status: Stable — non-breaking, clarifying revision of v2.5

This document is the authoritative specification of the .batest file format, an open, ZIP-based container format for storing reproducible blind audio comparison tests.

The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in RFC 2119.

v1.0 users: This document describes v2.6, a non-breaking revision of v2.0/v2.1/v2.2/v2.3/v2.4/v2.5 (formatVersion remains 2; see Top-Level Fields, testTypeConfig, swappedSetup, and Track Object Schema for what's new/clarified). Files produced by a v2.0, v2.1, v2.2, v2.3, v2.4, or v2.5 implementation remain fully valid under v2.6. The v1.0 specification (single-test test.json structure) remains available in this repository's git history via the v1.0 tag/release. See Migration from v1 for a summary of what changed since v1.

Overview

A .batest file is a ZIP container that fully describes one or more related blind audio comparison tests, grouped into a test set: the audio tracks being compared, the metadata needed to reproduce the listening experience, and optional supporting material (cover images, documentation, external references). A conforming .batest file is self-contained and MUST be playable offline, without any network access, by any implementation that supports the test type(s) it contains.

This specification describes the format independently of any single application or implementation. Test results (a listener's actual responses) are explicitly out of scope — see Test Results.

Design Goals

File Extension and Container Format

The format uses the .batest file extension. A .batest file MUST be a valid ZIP archive.

Directory Structure

MyTest.batest
│
├── manifest.json
├── testSet.json
├── required/
│   ├── track1.flac        (converted from WAV)
│   ├── track2.mp3         (lossy source, kept as-is)
│   └── track3.flac
├── assets/
│   ├── cover.jpg
│   ├── setup.jpg
│   ├── room.png
│   ├── measurements.pdf
│   └── intro.mp4
└── resources/

manifest.json

manifest.json is a lightweight, redundant summary of the test set. It allows a file browser or import dialog to display key information without unzipping and parsing testSet.json.

{
  "formatVersion": 2,
  "createdWith": "Blind Audio Test 1.0.0",
  "createdAt": "2026-07-10T18:30:00Z",
  "title": "Example Test",
  "comparisonCategory": "microphones",
  "models": [
    { "manufacturer": "Neumann", "model": "U87" },
    { "manufacturer": "AKG", "model": "C414" }
  ]
}

All fields listed above are REQUIRED.

v2 change: manifest.json no longer has a testType field. A test set can now contain multiple test[] objects, each with its own testType, so a single top-level value could no longer represent the test set. Determining the test type(s) in a test set now requires parsing testSet.json's test[] array.

testSet.json

testSet.json is the primary, authoritative description of the test set.

{
  "testSetId": "UUIDv7",
  "creatorUuid": null,
  "title": "Example Test",
  "description": "...",
  "comparisonCategory": "microphones",
  "content": {
    "type": "vocals",
    "genres": ["indie-pop"],
    "tags": ["soft", "breathy"]
  },
  "test": [
    {
      "testId": 0,
      "testType": "abx",
      "tracks": [],
      "backingTrack": {
        "file": "required/background.flac",
        "gainDb": -12
      },
      "loudnessMatching": {
        "mode": "integrated-lufs",
        "reference": "loudest",
        "version": 1
      },
      "trackLengthMode": "shortest",
      "testTypeConfig": {
        "abx": { "testRounds": 8 }
      }
    }
  ]
}

description and assets are omitted from the example above; see Schema Hygiene Convention and Assets for when each is present versus omitted.

Top-Level Fields

recording is not a valid field at the testSet.json root level in v2 — it is only meaningful at the test-object and track-object level. See Field Placement and Overrides.

v2.3 change: testSet.json gained a new OPTIONAL root-level field, poolAbxAcrossTests, allowing a test set's A/B/X test procedures to be evaluated together instead of separately. This is a non-breaking, additive change — formatVersion stays 2. The field is deliberately OPTIONAL rather than required, and its key is omitted entirely (never written as an explicit false) whenever pooling is not enabled, so that files written by a v2.0–v2.2 implementation (which omit this field) remain fully valid under v2.3 and are correctly interpreted as using separate, per-test evaluation — the only behavior that existed before this field was introduced. This is the first plain boolean field in this specification; unlike loudnessMatching (an object whose mere presence signals "enabled"), poolAbxAcrossTests carries its boolean value directly, but follows the same underlying convention: the default (false) is represented by omission, and only the non-default value (true) is ever written explicitly.

Test Object Schema

Each entry in testSet.json's test array describes one runnable test within the test set (for example, a test set could bundle an A/B/X test followed by a separate Rating test over the same source material).

{
  "testId": 0,
  "testType": "abx",
  "title": "Preamp A vs. Preamp B",
  "content": {},
  "recording": {},
  "soundSource": "vocals",
  "tracks": [],
  "backingTrack": null,
  "loudnessMatching": {
    "mode": "integrated-lufs",
    "reference": "loudest",
    "version": 1
  },
  "trackLengthMode": "shortest",
  "assets": {},
  "testTypeConfig": {
    "abx": { "testRounds": 8 }
  }
}

content, recording, and assets are shown above only to indicate where they are positioned in the object; each is OPTIONAL and omitted entirely when empty, per the Schema Hygiene Convention — see Field Placement and Overrides.

v2.1 change: each test object gained a new OPTIONAL title field, letting an individual test within a test set have its own display title, independent of the test set's root-level title. This is a non-breaking, additive change — formatVersion stays 2 — and files written by a v2.0 implementation (which omit this field) remain fully valid; implementations reading such a file fall back to a generated per-test label as before.

v2.2 change: each test object gained four new OPTIONAL fields — soundSource, soundSourceOther, soundSourceSubtype, and soundSourceSubtypeOther — identifying the kind of source material that individual test's tracks are drawn from, following the same raw-value-plus-free-text-fallback pattern already used by comparisonCategory/comparisonCategoryOther at the testSet.json root, but scoped per test object rather than per test set (see above for why). This is a non-breaking, additive change — formatVersion stays 2. Although the reference Blind Audio Test application currently always sets soundSource when creating a test, this specification deliberately keeps all four fields OPTIONAL (rather than adding soundSource to the test object's required fields): making a newly introduced field required at the schema level would retroactively invalidate every .batest file already produced under v2.0/v2.1, which is exactly what "non-breaking, additive" rules out. Files written by a v2.0 or v2.1 implementation (which omit all four fields) remain fully valid under v2.2, and implementations reading a test object without soundSource MUST treat it as "not specified" rather than erroring.

v2.5 change: each test object gained a new OPTIONAL swappedSetup field, letting an "ab" or "abx-then-ab" test object declare a second, position-swapped physical setup for its A/B comparison, to reduce the influence of position-dependent bias (e.g. swapped microphone assignment). See swappedSetup for the full field description. This is a non-breaking, additive change — formatVersion stays 2. Files written by a v2.0–v2.4 implementation (which omit this field) remain fully valid under v2.5; implementations reading a test object without swappedSetup MUST treat it as "no swapped setup recorded", the only state representable before this field existed.

Field Placement and Overrides

v2 introduces multiple valid nesting levels for content, recording, and assets. This is new relative to v1, where content and recording existed only at the (single) test level and the track level. Each field's valid levels are:

FieldtestSet.json roottest objecttrack object
content✅✅✅
recording❌✅✅
assets✅✅❌

All occurrences are OPTIONAL and, per the Schema Hygiene Convention, the key is omitted entirely at any level where there is nothing to add — never written as null or an empty object.

content and recording override rule: where the same field is present at more than one applicable level for a given track, the more specific level MUST take precedence over the less specific one, per field, not per object. For content, precedence from least to most specific is: testSet.json root → test object → track object. For recording, precedence is: test object → track object (there is no root-level recording). A consuming application MUST merge values field-by-field, with the most specific level's value winning when present, falling back to the next-less-specific level otherwise. Array-typed sub-fields (e.g. content.tags, content.genres, recording.signalChain) are NOT merged element-by-element: if a more specific level provides its own array for a given sub-field, that array replaces the less specific level's array for that sub-field entirely, rather than being concatenated with it.

Example. A microphone comparison test set with these fragments:

// testSet.json root
"content": { "type": "vocals", "tags": ["soft"] }
// test object (test[0])
"content": { "tags": ["breathy"] },
"recording": {
  "environment": "studio",
  "signalChain": [
    { "type": "preamp", "manufacturer": "Millennia", "model": "HV-3C" }
  ]
}
// track object (test[0].tracks[0])
"recording": {
  "signalChain": [
    { "type": "microphone", "manufacturer": "Neumann", "model": "U 87 Ai" }
  ]
}

For this track, the effective merged values are:

assets: unlike content and recording, this specification does not define an override or merge rule between testSet.json root-level assets and test-object-level assets — both, where present, are simply the assets relevant to that scope (test-set-wide documentation vs. test-specific documentation). How an implementation presents assets from both levels together is application-specific.

backingTrack

backingTrack is an OPTIONAL, single audio track that plays continuously in the background throughout a test (e.g. a full music production over which the individually compared elements — instrument takes, effects, etc. — are layered). It is not part of the blind comparison itself: it plays unchanged, in parallel, for every track being evaluated in that test.

This field MUST always be present on a test object; it is null when that test does not use a backing track (see Schema Hygiene Convention).

"backingTrack": {
  "file": "required/background.flac",
  "gainDb": -12
}

loudnessMatching

loudnessMatching describes the loudness-matching behavior used for a test. The key is OPTIONAL and is omitted entirely when loudness matching was not enabled for that test — there is no enabled: false representation; the block's mere presence on the test object means it was enabled.

"loudnessMatching": {
  "mode": "integrated-lufs",
  "reference": "loudest",
  "version": 1
}

Implementations MUST accept files that predate this convention and may still contain a legacy enabled and/or target field inside the block; both MUST be ignored on import, since the block's presence already conveys "enabled", and no known implementation ever wrote a target value other than matching relative to reference.

trackLengthMode

trackLengthMode determines how tracks with different durations are handled during playback of a test:

This field is OPTIONAL and is only meaningful — and only ever written — when that test's tracks actually have different durations; it is omitted entirely when all of the test's tracks have the same length (or there is only one track). Implementations MUST assume "shortest" when the key is absent but the tracks do have different lengths.

trackLengthMode does not apply to a test object whose testType is rating (testTypeConfig.rating): a Rating test presents one track at a time rather than playing multiple tracks in sync, so this field is never written for that test type, regardless of whether the tracks have different durations. Implementations MUST ignore this key if present on a Rating test object (e.g. from a file exported by an older implementation).

Track Object Schema

Each entry in a test object's tracks[] follows this schema, independent of test type, so the same schema is used for A/B, A/B/X, and multitrack Ranking tests alike.

{
  "trackId": 0,
  "filename": "track1.flac",
  "originalFilename": "Vocal.wav",
  "manufacturer": "Neumann",
  "model": "U87",
  "label": null,
  "recording": {
    "signalChain": [
      { "type": "microphone", "manufacturer": "Neumann", "model": "U 87 Ai" }
    ]
  },
  "originalFormat": "wav",
  "storedFormat": "flac",
  "originalSampleRate": 48000,
  "originalBitDepth": 24,
  "durationSeconds": 187.4,
  "integratedLufs": -18.7
}

Example — a track with a free-text "Others" manufacturer/model:

{
  "trackId": 0,
  "filename": "track1.flac",
  "originalFilename": "Vocal.wav",
  "manufacturer": "others",
  "model": "others",
  "manufacturerOther": "Acme Custom Mic Co.",
  "modelOther": "Prototype X1"
}

A consuming implementation that wants a single display string for this track falls back to manufacturerOther/modelOther whenever manufacturer/model equal "others" — the same pattern already used for comparisonCategory when producing manifest.json's resolved comparisonCategory value (and how this track's data ends up represented, already resolved, in manifest.json's models array).

v2.6 change: clarified that a track object's manufacturer / model fields MUST hold the stable, human-readable name of the item under comparison rather than a database ID, internal slug, or other instance-dependent identifier, and added minLength: 1 to both fields in schema/testSet.schema.json. This is a non-breaking clarification of existing, previously underspecified behavior — formatVersion stays 2, and no field was added, removed, or changed in shape. Files already correctly writing human-readable names for these fields are unaffected. Files from an implementation that (incorrectly, per this clarification) wrote a raw ID into these fields remain structurally valid .batest files under this specification's schema either way — the affected consuming application handles any resulting display fallback on its own; no migration or backfill of previously written .batest files is defined or required by this specification.

Format Conversion Rule

testTypeConfig

Test-type-specific data lives in a namespaced object, keyed by test type, so the fields shared across the format (tracks, loudnessMatching, playback, randomization, assets, resources) remain identical across all test types, and new test types can be added without breaking older parsers — unknown testTypeConfig keys MUST be ignored by implementations that do not support them. The object's single key MUST match the test object's own testType field (see Test Object Schema).

This specification defines the following test types: A/B, A/B/X, A/B/X→A/B, Ranking, and Rating.

A/B:

"testType": "ab",
"testTypeConfig": {
  "ab": {
    "testRounds": 2
  }
}

A/B/X:

"testType": "abx",
"testTypeConfig": {
  "abx": {
    "testRounds": 8
  }
}

A/B/X→A/B:

This combined procedure runs two phases back to back, each with its own independent round count, so it uses two fields instead of a single testRounds:

"testType": "abx-then-ab",
"testTypeConfig": {
  "abx-then-ab": {
    "abxRounds": 8,
    "abRounds": 2,
    "requiresSuccessfulAbx": true
  }
}

requiresSuccessfulAbx is an OPTIONAL boolean, scoped to this abx-then-ab procedure. When true, the A/B phase only runs after the listener completed the preceding A/B/X phase successfully; when false, the A/B phase always runs, regardless of the A/B/X outcome — this was the only behavior available before this field existed. The key is omitted entirely when false, per the Schema Hygiene Convention; files written by a v2.0–v2.3 implementation (which omit this field) remain fully valid under v2.4 and are correctly interpreted as "always run". (New in v2.4 — see the note below.)

Ranking:

"testType": "ranking",
"testTypeConfig": {
  "ranking": {
    "testRounds": 3
  }
}

For ab, abx, and ranking, testRounds is the number of rounds the listener completes. For abx-then-ab, abxRounds and abRounds are the round counts for the A/B/X phase and the following A/B phase, respectively. All fields shown above are REQUIRED for their respective test type.

Rating:

"testType": "rating",
"testTypeConfig": {
  "rating": {
    "ratingMode": "identified",
    "ratingCategories": [
      {
        "id": 0,
        "label": "Brightness",
        "scaleType": "numeric",
        "scaleMin": 1,
        "scaleMax": 5,
        "scaleMinLabel": "Dull",
        "scaleMaxLabel": "Bright"
      },
      {
        "id": 1,
        "label": "Low-End",
        "scaleType": "bipolar",
        "scaleRadius": 5,
        "centerLabel": "Just right",
        "negativeLabel": "Too little",
        "positiveLabel": "Too much"
      }
    ]
  }
}

v2.4 change: testTypeConfig's abx-then-ab object gained a new OPTIONAL field, requiresSuccessfulAbx, letting an A/B/X→A/B procedure be configured so the A/B phase is gated on a successful preceding A/B/X phase, instead of always running unconditionally. This is a non-breaking, additive change — formatVersion stays 2. The field's key is omitted entirely (never written as an explicit false) whenever it is not enabled, so that files written by a v2.0–v2.3 implementation (which omit this field) remain fully valid under v2.4 and are correctly interpreted as "always run" — the only behavior that existed before this field was introduced.

Additional test types, such as MUSHRA, are under consideration for a future version of this specification but are not yet defined. See the project README for status.

swappedSetup

swappedSetup is an OPTIONAL field on a test object, describing a second, position-swapped physical recording/reproduction setup for that test's A/B comparison, used to reduce the influence of a position-dependent bias that has nothing to do with the actual items being compared. A common example is a microphone shootout where two microphones each feed a different preamp: if Mic 1 always feeds Preamp A and Mic 2 always feeds Preamp B, a listener's judgment of "Preamp A vs. Preamp B" is confounded with "Mic 1 vs. Mic 2" and, more generally, with the microphones' physical positions in the room. Recording a second pass with the microphones' positions swapped (Mic 2 → Preamp A, Mic 1 → Preamp B) and presenting both passes lets a listener's judgment be checked for consistency independent of which physical position fed which compared item.

{
  "testId": 0,
  "testType": "ab",
  "tracks": [
    { "trackId": 0, "filename": "track1.flac", "...": "Preamp A, Mic 1" },
    { "trackId": 1, "filename": "track2.flac", "...": "Preamp B, Mic 2" }
  ],
  "swappedSetup": {
    "tracks": [
      { "trackId": 0, "filename": "track1b.flac", "...": "Preamp A, Mic 2" },
      { "trackId": 1, "filename": "track2b.flac", "...": "Preamp B, Mic 1" }
    ]
  }
}

The following rules govern swappedSetup:

This specification deliberately does not define a mechanically validated (JSON Schema) enforcement of the 1:1 track-count coupling or the manufacturer/model matching rule above — both are normative prose requirements only, consistent with how this specification already treats the positional trackId/testId rules elsewhere (see Identifier Stability).

Assets

Assets are optional local files that help explain or document the test set but never change the listening test itself.

Supported asset types: Images (JPG, PNG, WebP), Videos (MP4 recommended), PDF documents.

testSet.json references assets as follows, at both the root level and, optionally, at the level of an individual test object (see Field Placement and Overrides):

"assets": {
  "cover": "assets/cover.jpg",
  "items": [
    { "type": "image", "file": "assets/setup.jpg", "label": "Microphone setup" },
    { "type": "video", "file": "assets/intro.mp4", "label": "Introduction" },
    { "type": "pdf", "file": "assets/measurements.pdf", "label": "Measurement report" }
  ]
}

The assets key itself is OPTIONAL and is omitted entirely from wherever it would otherwise appear in testSet.json when neither cover nor items has actual content. When assets is present, it MUST contain only the sub-field(s) that actually have content — cover alone, items alone, or both — never an empty placeholder for the field that has no content. For example, a test set with only a cover image but no additional assets exports as:

"assets": {
  "cover": "assets/cover.jpg"
}

Missing assets MUST NOT prevent the test from running.

Resources

resources is an OPTIONAL array of external references that are not embedded in the container (e.g. YouTube videos, manufacturer pages, whitepapers, AES papers, Git repositories). The key is omitted entirely when empty.

"resources": [
  { "type": "link", "title": "Manufacturer page", "url": "https://..." },
  { "type": "paper", "title": "AES Convention Paper 12345", "url": "https://..." }
]

Test Results

Results are never stored inside .batest. A test's results are linked to the test set by its testSetId (see Identifier Stability) in whatever backend or storage system an implementation uses; this specification does not define a results format.

Data Integrity and Security

File Integrity

File integrity is covered by the ZIP container's own built-in per-entry CRC32 checksum, which detects corruption during storage or transfer. Implementations SHOULD rely on this mechanism, together with standard ZIP validation, when reading a .batest file.

Zip-Slip Protection

When extracting a .batest container, implementers MUST validate every entry path before writing it to disk. Entries containing ../, absolute paths, or that otherwise resolve outside the intended extraction directory MUST be rejected. This is a standard ZIP extraction vulnerability known as "Zip Slip", and applies regardless of file extension.

Canonical Serialization for Hashing

Implementers who need to compute a hash over the content of testSet.json or manifest.json — for example, to detect whether the payload has changed, for caching, or for synchronization between systems — SHOULD compute that hash over the canonical serialization of the JSON content, as defined by RFC 8785, the JSON Canonicalization Scheme (JCS), rather than over the raw file bytes.

Different JSON serializers can produce different key ordering, whitespace, or number formatting for otherwise identical data. Hashing raw bytes would therefore cause semantically identical content to produce different hashes depending on which tool or library wrote the file. Canonicalizing first removes this variation. JCS implementations already exist for most common languages, including JavaScript/ TypeScript, Python, and Rust, so implementers do not need to write their own canonicalizer.

This recommendation applies to any tooling built around the format (validators, sync tools, caches, etc.); it is not a requirement for the container format itself, and does not affect how testSet.json or manifest.json are written to disk inside a .batest file.

Schema Hygiene Convention

Optional fields in manifest.json and testSet.json follow a consistent convention: when an optional field has no meaningful value, its key is omitted from the JSON entirely — it is never written as null, an empty object {}, or an empty array []. What counts as "empty" depends on the field's type: for object-typed fields it means an empty object or no meaningful sub-fields set; for array-typed fields it means an empty array; for string-typed fields it means no value present; for number-typed fields it means no measurement or value available; for boolean-typed fields it means the field's documented default value (so far, always false) — only the non-default value is ever written explicitly, and the key is omitted rather than writing the default out.

Implementations MUST accept both the omitted-key form and an explicit null (or empty {}/[]) for these fields, since files exported by older implementations may still contain the legacy representation.

This convention applies to (among others): description, assets (and its cover and items sub-fields, at whichever level assets appears), playback, randomization, resources, content, recording (at whichever level each appears — testSet root, test object, or track object, per Field Placement and Overrides), comparisonCategoryOther, comparisonSubcategory, comparisonSubcategoryOther, loudnessMatching, trackLengthMode, each test object's title (new in v2.1), each test object's soundSource, soundSourceOther, soundSourceSubtype, and soundSourceSubtypeOther (new in v2.2), poolAbxAcrossTests (testSet.json root, new in v2.3), requiresSuccessfulAbx (testTypeConfig.abx-then-ab, new in v2.4), swappedSetup (each test object, new in v2.5), and, on each track object, manufacturer, model, manufacturerOther, modelOther, notes, and integratedLufs.

A small number of fields are exceptions and are instead always present with an explicit null value when unset, rather than being omitted: creatorUuid (testSet.json root), backingTrack (on each test object), label (on each track object), and scaleMinLabel/scaleMaxLabel (on numeric rating categories). originalBitDepth (on each track object) is likewise always present, but is null for a different reason: not because the value is unset, but because the concept of bit depth does not apply to lossy-compressed formats.

Identifier Stability

testSet.json's top-level testSetId field is the test set's unique, permanent identifier (a UUIDv7). This identifier MUST be treated as stable and immutable: once assigned to a test set, it MUST NOT change, even across subsequent edits to any other field in testSet.json. Systems that reference a test set externally (e.g. to associate results with it, per Test Results) rely on testSetId remaining constant for the lifetime of the test set.

This is unrelated to the per-test testId and per-track trackId fields described above, both of which are positional indices scoped to their containing array (test and a given test object's tracks, respectively) rather than standalone stable identifiers — they renumber if their containing array is reordered, unlike testSetId.

Versioning and Compatibility

Migration from v1

v2 (formatVersion: 2) is a breaking revision of v1 (formatVersion: 1). A v1 test.json file is not structurally compatible with v2's testSet.json and requires conversion; this is a conceptual summary of what changed, not a full migration guide:

No automatic v1-to-v2 conversion tool is defined by this specification.

Guiding Principle

A .batest file represents one or more fully reproducible listening tests. Required assets make the test set executable. Optional assets (including embedded videos) provide additional context. External resources complement the package without increasing its size unnecessarily.