Testcontainers Test Data: Seed Integration Tests With Real Fixtures
Testcontainers test data is the half of the problem Testcontainers
leaves to you. The library answers "which database do my integration
tests run against": a real, throwaway Postgres, MySQL, or MongoDB in
Docker, started for the test run and removed afterwards. It doesn't
answer "what's in it". An empty container proves only that your queries
parse, and the usual fix, a few hand-written INSERT statements, gives
you three rows that drifted from the schema two migrations ago and have
no relations worth testing.
This post covers the other half: generating realistic, related seed data for integration tests from one template with a fixed seed, committing it as JSON, and loading it into the container from the test's setup hook. The examples use PostgreSQL with JUnit 5, Vitest, and pytest. The pattern is the same for any database your driver can insert into.
JsonFabrica's part is narrow. It's a JSON-over-HTTP generation API with no Testcontainers module, no database connector, and no Docker image. It never starts a container or writes a row. Your test code reads the JSON and inserts it with the driver or ORM you already use.
Why an empty Testcontainers database proves nothing
A repository method like "find overdue invoices" has several branches
hiding in its WHERE clause: drafts that were never issued, voided
invoices, paid ones, open ones due before and after the cut-off date.
A test against an empty table passes for every one of those branches,
including the wrong ones. A test against three hand-written rows covers
whichever branches the author thought of that day.
Hand-written integration test fixtures also fail in quieter ways:
- They drift from the schema. A column becomes
NOT NULLor gets renamed, and the fixture SQL is the last thing anyone updates. - They lack relations. Writing consistent foreign keys by hand is tedious, so fixtures tend to hold one customer with one invoice, and the query that aggregates per customer never sees more than one row.
- They skip empty states. A customer with no invoices at all is the
row that catches a
LEFT JOINthat should have been anINNER JOINor aSUMthat returnsNULLinstead of 0, and it's the row nobody types.
Seeding a long-lived local dev database is a different job, with different row counts and refresh rules. The local dev seeding guide covers that. Browser tests against a full running backend have their own setup, covered in Playwright test data. Here the database lives for one test class or one test run, and the fixture has to be small, exact, and the same every time.
Generate Testcontainers test data once, with a fixed seed
The example schema is a small billing database: customers, and invoices that belong to them.
-- db/schema.sql
CREATE TABLE customers (
id bigint GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
name text NOT NULL,
email text NOT NULL UNIQUE,
country char(2) NOT NULL,
created_at timestamptz NOT NULL
);
CREATE TABLE invoices (
id bigint GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
customer_id bigint NOT NULL REFERENCES customers (id),
number text NOT NULL UNIQUE,
status text NOT NULL
CHECK (status IN ('draft', 'open', 'paid', 'void')),
amount_cents integer NOT NULL CHECK (amount_cents >= 0),
currency char(3) NOT NULL,
issued_at timestamptz,
due_at timestamptz,
paid_at timestamptz
);
One template produces both tables as one JSON document, with keys that match the column names:
{
"customers": [
<for(i, 1, getParam('customers', 20))><if(getVar('i') > 1)>,<endIf>
<setVar('first', getRandomName())><setVar('last', getRandomSurname())>
{
"id": <getVar('i')>,
"name": "<getVar('first')> <getVar('last')>",
"email":
"<toLowerCase(getVar('first'))>.<toLowerCase(getVar('last'))>.<getVar('i')>@example.com",
"country": "<getRandomElement('DE', 'FR', 'GB', 'NL', 'US')>",
"created_at": "<getRandomDate('2025-01-01', '2026-06-30')>"
}<end_for>
],
"invoices": [
<for(i, 1, getParam('invoices', 60))><if(getVar('i') > 1)>,<endIf>
<if(getVar('i') == 1)><setVar('status', 'draft')>
<elseIf(getVar('i') == 2)><setVar('status', 'void')>
<else><setVar('status', getRandomElement('paid', 'paid', 'paid', 'open'))>
<endIf>
{
"id": <getVar('i')>,
"customer_id": <getRandomNumber(2, getParam('customers', 20))>,
"number": "INV-2026-<appendBefore('0', getVar('i'), 5)>",
"status": "<getVar('status')>",
"amount_cents": <getRandomNumber(900, 250000)>,
"currency": "EUR",
"issued_at": <if(getVar('status') == 'draft')>null
<else>"<getRandomDate('2026-07-01', '2026-08-31')>"<endIf>,
"due_at": <if(getVar('status') == 'draft')>null
<else>"<getRandomDate('2026-09-01', '2026-09-30')>"<endIf>,
"paid_at": <if(getVar('status') == 'paid')>
"<getRandomDate('2026-09-01', '2026-09-30')>"<else>null<endIf>
}<end_for>
]
}
How it works:
- Row counts are parameters.
getParamreadscustomersandinvoicesfrom the request, with defaults of 20 and 60. Loop bounds are inclusive, and the comma guard<if(getVar('i') > 1)>,<endIf>puts a comma before every row except the first. The control flow docs coverforandif/elseIf/else. - Foreign keys are valid by construction. Customer ids are the loop
index, 1 to 20. Each invoice picks its
customer_idwithgetRandomNumber, whose bounds are inclusive, over the range 2 to 20, so every invoice points at a customer that exists. Most customers get several invoices, and a few get none by chance (with seed20261007, customers 6 and 17). - Customer 1 never has an invoice. Starting the range at 2 plants the empty state on purpose, so there's always one known customer with nothing to sum.
- The first invoices are planted edge cases. Invoice 1 is a
draftwith noissued_atordue_at, and invoice 2 isvoid. The rest arepaidthree times as often asopen, becausegetRandomElementpicks evenly from its arguments andpaidis listed three times. - Unique columns are unique by construction.
emailhas aUNIQUEconstraint, and two customers can draw the same name, so the loop index goes into the address. Don't count on random values to satisfy a unique index. - Variables only where a value is used twice.
firstandlastbuild both the name and the email.statusis printed and decides which dates are set. Everything else is written inline in its field. - Dates stay in order without arithmetic. Template expressions
support comparisons and
&&/||but no arithmetic, so "issued plus 30 days" isn't available. Non-overlapping ranges forgetRandomDatedo the same job: customers sign up before July 2026, invoices are issued in July or August, and due and payment dates fall in September.
The long email line stays long. A line break in the literal text of a
quoted value ends up in the output. Breaks are only safe inside a
placeholder's parentheses, which wouldn't make this line any easier to
read.
Save the template as scripts/billing.tmpl and render it with a short
script that pins the seed. It sends the template to
POST /v1/templates/generate, checks
that the seed was applied, and writes only the generated document:
#!/usr/bin/env bash
# scripts/billing-fixture.sh: regenerate the integration test fixture.
set -euo pipefail
SEED=20261007 # change only when you mean to replace the data
OUT=src/test/resources/fixtures/billing.json
TMP=$(mktemp)
trap 'rm -f "$TMP"' EXIT
jq -n --rawfile body scripts/billing.tmpl --argjson seed "$SEED" \
'{body: $body, seed: $seed, params: {customers: 20, invoices: 60}}' \
| curl -sS --fail -X POST https://api.jsonfabrica.com/v1/templates/generate \
-H "Authorization: Bearer $JSONFABRICA_API_KEY" \
-H "Content-Type: application/json" \
-d @- -o "$TMP"
# Stop if the seed wasn't applied: nobody could regenerate the file.
jq -e --argjson s "$SEED" '.meta.seed == $s' "$TMP" > /dev/null
mkdir -p "$(dirname "$OUT")"
jq '.data' "$TMP" > "$OUT"
Set OUT to wherever your test runner reads fixtures. For Java that's
the test classpath, as shown, and schema.sql belongs there too, as
src/test/resources/db/schema.sql. The Node example reads
test/fixtures/billing.json and the Python one
tests/fixtures/billing.json; both read db/schema.sql from the
project root.
The generate response has two fields: data, the generated document,
and meta, whose seed echoes the seed that was used. Commit the JSON
file next to the test code. CI then needs no network access to
JsonFabrica and no API key, and the data changes only when someone
reruns the script and reviews the diff.
Rendered with seed 20261007, the fixture starts like this:
{
"customers": [
{
"id": 1,
"name": "Gilander Bacon",
"email": "gilander.bacon.1@example.com",
"country": "DE",
"created_at": "2025-11-23T01:44:31.394Z"
}
],
"invoices": [
{
"id": 1,
"customer_id": 7,
"number": "INV-2026-00001",
"status": "draft",
"amount_cents": 239078,
"currency": "EUR",
"issued_at": null,
"due_at": null,
"paid_at": null
},
{
"id": 2,
"customer_id": 7,
"number": "INV-2026-00002",
"status": "void",
"amount_cents": 62035,
"currency": "EUR",
"issued_at": "2026-07-03T10:25:04.538Z",
"due_at": "2026-09-23T10:27:14.532Z",
"paid_at": null
}
]
}
The arrays continue to 20 customers and 60 invoices. Customer 7 ended up
with the draft, the void invoice, one open invoice, INV-2026-00003, of
140,663 cents, and two paid ones, which makes it a useful customer for
the balance tests below.
Testcontainers database seeding in one SQL statement per table
The loader doesn't need a JSON library or an ORM. PostgreSQL's
json_populate_recordset turns a JSON array into rows of a table's own
type, matching object keys to column names, so each table is one
INSERT ... SELECT:
INSERT INTO customers
SELECT * FROM json_populate_recordset(
null::customers, $1::json -> 'customers'
);
Three details make this a reliable loader:
- Parents before children. Insert
customersbeforeinvoices, or the foreign key rejects the first invoice. - Advance the identity sequences. The fixture supplies explicit ids,
which a
GENERATED BY DEFAULT AS IDENTITYcolumn accepts without advancing its sequence. The first row a test inserts without an id would get id 1 and fail with a duplicate key error. Onesetvalper table moves the sequence past the fixture's highest id. - Schema drift fails loudly where it matters. A JSON key that no
longer matches a column is ignored and the column gets
NULL. For aNOT NULLcolumn, the insert fails with an error naming the column, so a renamed column surfaces on the next test run. Nullable columns don't get that protection, so regenerate the fixture whenever a migration touches these tables.
For MySQL, JSON_TABLE plays the same role. For other databases, parse
the file and batch-insert with your driver or ORM. The
ORM seed script alternative post
shows createMany and bulk_create versions of that.
Here is the loader in Java, using plain JDBC:
package com.example.billing;
import java.io.InputStream;
import java.nio.charset.StandardCharsets;
import java.sql.Connection;
import java.sql.PreparedStatement;
import java.sql.Statement;
import java.util.List;
final class FixtureLoader {
// Parents before children, so foreign keys resolve on insert.
private static final List<String> TABLES =
List.of("customers", "invoices");
static void load(Connection c, String resource) throws Exception {
String json;
try (InputStream in = FixtureLoader.class
.getResourceAsStream(resource)) {
json = new String(in.readAllBytes(), StandardCharsets.UTF_8);
}
for (String t : TABLES) {
try (PreparedStatement insert = c.prepareStatement(
"INSERT INTO " + t + " SELECT * FROM json_populate_recordset("
+ "null::" + t + ", ?::json -> '" + t + "')")) {
insert.setString(1, json);
insert.executeUpdate();
}
try (Statement st = c.createStatement()) {
st.execute("SELECT setval(pg_get_serial_sequence('" + t
+ "', 'id'), (SELECT max(id) FROM " + t + "))");
}
}
}
}
The table names are concatenated into the SQL because they come from a constant list, never from input. The fixture itself always goes in as a bind parameter.
JUnit 5: load seed data for integration tests in @BeforeAll
With Testcontainers for Java, the @Testcontainers annotation enables
the JUnit 5 extension, and every field annotated with @Container gets
started and stopped for you. The field's modifier decides the
lifecycle:
- A static
@Containerfield is started once before the first test method in the class and stopped after the last one. - An instance
@Containerfield is started and stopped around every test method.
The dependencies, at the time of writing, are
org.testcontainers:testcontainers-junit-jupiter:2.0.5,
org.testcontainers:testcontainers-postgresql:2.0.5, and the
PostgreSQL JDBC driver (org.postgresql:postgresql), which the
Testcontainers module doesn't pull in for you.
package com.example.billing;
import static org.junit.jupiter.api.Assertions.assertEquals;
import java.sql.Connection;
import java.sql.DriverManager;
import java.time.Instant;
import java.util.List;
import org.junit.jupiter.api.AfterAll;
import org.junit.jupiter.api.BeforeAll;
import org.junit.jupiter.api.Test;
import org.testcontainers.junit.jupiter.Container;
import org.testcontainers.junit.jupiter.Testcontainers;
import org.testcontainers.postgresql.PostgreSQLContainer;
@Testcontainers
class InvoiceRepositoryTest {
@Container
static PostgreSQLContainer postgres =
new PostgreSQLContainer("postgres:17-alpine")
.withInitScript("db/schema.sql");
static Connection conn;
static InvoiceRepository repo; // stands in for your code under test
@BeforeAll
static void loadFixture() throws Exception {
conn = DriverManager.getConnection(postgres.getJdbcUrl(),
postgres.getUsername(), postgres.getPassword());
FixtureLoader.load(conn, "/fixtures/billing.json");
repo = new InvoiceRepository(conn);
}
@AfterAll
static void close() throws Exception {
conn.close();
}
@Test
void findsOverdueInvoices() {
List<String> overdue =
repo.findOverdue(Instant.parse("2026-09-15T00:00:00Z"));
assertEquals(List.of("INV-2026-00006", "INV-2026-00011",
"INV-2026-00043", "INV-2026-00048"), overdue);
}
@Test
void openBalanceIgnoresDraftAndVoidInvoices() {
assertEquals(140_663, repo.openBalanceCents(7));
}
@Test
void customerWithoutInvoicesHasZeroBalance() {
assertEquals(0, repo.openBalanceCents(1));
}
}
The order of events for this class:
- The extension starts the static container.
withInitScriptrunsdb/schema.sqlfrom the test classpath after the database is up and before any test code gets a connection. @BeforeAllruns after the extension's own setup, so the container is already running. It connects and loads the fixture.- The three tests run against the same 80 rows.
- After the last test, the extension stops the container and the data is gone.
The assertions are exact because the fixture is committed. With seed
20261007, four open invoices are due before September 15. Customer
7's balance is the one open invoice, 140,663 cents. A query that
forgets to exclude the draft or the void invoice returns a larger
number and fails. Customer 1 has no invoices, so the balance query has
to return 0 rather than NULL.
withInitScript is fine for an example, but a hand-kept schema.sql
can drift from production just like hand-written inserts. In a real
project, run your migrations against the container in @BeforeAll
before loading the fixture, for example with
Flyway.configure().dataSource(url, user, password).load().migrate().
Data that exercises the migrations themselves is a separate job,
covered in test data for database migrations.
On Testcontainers for Java 1.x, the class is
org.testcontainers.containers.PostgreSQLContainer, it's generic
(PostgreSQLContainer<?>), and the artifacts are named junit-jupiter
and postgresql without the testcontainers- prefix. The lifecycle is
the same.
Per-class vs per-suite containers, and where data loading goes
A static @Container gives every test class its own database with a
fresh copy of the fixture. That's the simplest isolation there is: no
class can see another class's writes. The cost is one container start
per class, typically a few seconds each, which adds up across dozens of
classes.
The alternative is the singleton container pattern from the
Testcontainers docs: start the container in a static initializer of a
base class, without @Container, and let every test class extend it.
The container starts once per JVM, and the Ryuk container that
Testcontainers starts alongside your tests removes it when the JVM
exits. Load the fixture in the same static block, so it also happens
once:
// Imports as in the test above, plus java.sql.SQLException.
abstract class PostgresTestBase {
static final PostgreSQLContainer POSTGRES =
new PostgreSQLContainer("postgres:17-alpine")
.withInitScript("db/schema.sql");
static {
POSTGRES.start();
try (Connection c = connect()) {
FixtureLoader.load(c, "/fixtures/billing.json");
} catch (Exception e) {
throw new ExceptionInInitializerError(e);
}
}
static Connection connect() throws SQLException {
return DriverManager.getConnection(POSTGRES.getJdbcUrl(),
POSTGRES.getUsername(), POSTGRES.getPassword());
}
}
Now every class shares one copy of the data, so a test that writes has to clean up after itself. The cheapest way is a transaction per test that's always rolled back:
// Also imports org.junit.jupiter.api.BeforeEach and AfterEach.
class InvoiceWriteTest extends PostgresTestBase {
Connection tx;
@BeforeEach
void begin() throws SQLException {
tx = connect();
tx.setAutoCommit(false);
}
@AfterEach
void rollback() throws SQLException {
tx.rollback();
tx.close();
}
@Test
void voidingTheOpenInvoiceClearsTheBalance() {
InvoiceRepository repo = new InvoiceRepository(tx);
repo.voidInvoice("INV-2026-00003");
assertEquals(0, repo.openBalanceCents(7));
}
}
This only works if the code under test uses the connection you hand it
and doesn't commit on its own. Spring's test transactions follow the
same idea. When the code manages its own transactions, truncate and
reload instead. Reloading 80 rows is two statements per table, so
running TRUNCATE invoices, customers RESTART IDENTITY followed by
FixtureLoader.load in @BeforeEach is cheap. In Node, the
@testcontainers/postgresql module also has snapshot() and
restoreSnapshot(): take a snapshot right after loading the fixture and
restore it between tests. Close your connections before either call,
because the module drops and recreates the database.
Rule of thumb: per-class containers while the suite is small or tests write a lot, a singleton once container startup dominates the run.
Testcontainers with Vitest or Jest: the Node lifecycle
The Node library follows the same model. Install its PostgreSQL module and a driver:
npm install --save-dev @testcontainers/postgresql pg @types/pg
The loader is the same SQL, using pg. It takes the JSON as a string
so it works with a committed file or a fetched one:
// test/load-fixture.ts
import type { Client } from "pg";
// Parents before children, so foreign keys resolve on insert.
const TABLES = ["customers", "invoices"] as const;
export async function loadFixture(db: Client, json: string) {
for (const t of TABLES) {
await db.query(
`INSERT INTO ${t} SELECT * FROM ` +
`json_populate_recordset(null::${t}, $1::json -> '${t}')`,
[json],
);
await db.query(
`SELECT setval(pg_get_serial_sequence('${t}', 'id'), ` +
`(SELECT max(id) FROM ${t}))`,
);
}
}
A test file starts its own container in beforeAll, creates the
schema, loads the fixture, and stops the container in afterAll:
// test/invoices.test.ts
import { readFile } from "node:fs/promises";
import {
PostgreSqlContainer,
type StartedPostgreSqlContainer,
} from "@testcontainers/postgresql";
import { Client } from "pg";
import { afterAll, beforeAll, expect, test } from "vitest";
import { InvoiceRepository } from "../src/invoice-repository";
import { loadFixture } from "./load-fixture";
let container: StartedPostgreSqlContainer;
let db: Client;
beforeAll(async () => {
container = await new PostgreSqlContainer("postgres:17-alpine")
.start();
db = new Client({ connectionString: container.getConnectionUri() });
await db.connect();
await db.query(await readFile("db/schema.sql", "utf8"));
await loadFixture(db, await readFile("test/fixtures/billing.json",
"utf8"));
}, 120_000); // the first run may pull the image
afterAll(async () => {
await db?.end();
await container?.stop();
});
test("finds invoices overdue on September 15", async () => {
const repo = new InvoiceRepository(db);
const overdue = await repo.findOverdue(new Date("2026-09-15"));
expect(overdue).toEqual([
"INV-2026-00006",
"INV-2026-00011",
"INV-2026-00043",
"INV-2026-00048",
]);
});
test("a customer with no invoices has a zero balance", async () => {
expect(await new InvoiceRepository(db).openBalanceCents(1)).toBe(0);
});
pg runs a query string without parameters as a simple query, so the
multi-statement schema.sql goes through in one call. The 120-second
timeout matters: Vitest's default hook timeout is 10 seconds, which an
image pull on a cold CI runner can exceed. Jest's beforeAll takes the
same timeout argument, and the rest of the file is unchanged apart from
the imports. The image argument to PostgreSqlContainer is required in
recent versions of the module, so always pass a pinned tag.
For one container per run instead of one per file, start it in a Vitest
globalSetup file and pass the connection URI to the tests with
project.provide and inject. That's the Node equivalent of the Java
singleton, and the same rule applies: tests that write must roll back
or reload.
Python: the same idea as a pytest fixture
testcontainers-python fits the same lifecycle into a pytest fixture.
The fixture's scope plays the role of the static or instance field:
scope="module" gives each test module its own database, and
scope="session" shares one across the run.
# tests/conftest.py
from pathlib import Path
import psycopg
import pytest
from testcontainers.community.postgres import PostgresContainer
TABLES = ("customers", "invoices") # parents before children
@pytest.fixture(scope="module")
def db():
with PostgresContainer("postgres:17-alpine", driver=None) as pg:
with psycopg.connect(pg.get_connection_url()) as conn:
conn.execute(Path("db/schema.sql").read_text())
data = Path("tests/fixtures/billing.json").read_text()
for t in TABLES:
conn.execute(
f"INSERT INTO {t} SELECT * FROM "
f"json_populate_recordset(null::{t}, "
f"%s::json -> '{t}')",
(data,),
)
conn.execute(
f"SELECT setval(pg_get_serial_sequence('{t}', 'id'), "
f"(SELECT max(id) FROM {t}))"
)
conn.commit()
yield conn
driver=None makes get_connection_url return a plain postgresql://
URL that psycopg accepts, instead of an SQLAlchemy URL with
+psycopg2. Recent releases import PostgresContainer from
testcontainers.community.postgres, and older ones from
testcontainers.postgres.
Parallel test workers and data isolation
Parallel workers are where per-container isolation pays off, and the mechanism is the container, not the data:
- Vitest and Jest run test files in separate worker processes, and
Vitest runs files in parallel by default. A container started in a
file's
beforeAllbelongs to that file. Ten files running at once means ten databases, each with its own copy of the same fixture, so identical ids in each never collide. - JUnit 5: the Testcontainers extension is documented as tested
only with sequential execution, and parallel use inside one JVM is
unsupported. Parallelize with JVM forks instead, such as Maven
Surefire's
forkCountor Gradle'smaxParallelForks. Each fork loads its own classes, so static containers and singletons are per fork.
Testcontainers maps each container port to a random free host port, so parallel containers don't fight over 5432. The only shared resource is the Docker host, and that's the limit to watch: a CI runner with two CPUs won't run twelve Postgres containers well. A suite-wide container shared by parallel workers needs a separate database or schema per worker, because the workers do see each other's rows.
Fetch integration test fixtures in setup instead of committing them
A committed file is the right default. Fetching in setup is useful when you want a fresh, still reproducible dataset on some runs, such as a nightly job that tries a new seed each night to shake out data the fixed fixture never contains.
Save the template once with POST /v1/templates, then call
POST /v1/templates/{templateId}/generate from the test setup:
// test/fetch-fixture.ts
export async function fetchFixture(): Promise<string> {
const seed = process.env.FIXTURE_SEED; // unset: the API picks one
const res = await fetch(
"https://api.jsonfabrica.com/v1/templates/" +
`${process.env.FIXTURE_TEMPLATE_ID}/generate`,
{
method: "POST",
headers: {
Authorization: `Bearer ${process.env.JSONFABRICA_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
params: { customers: 20, invoices: 60 },
...(seed ? { seed: Number(seed) } : {}),
}),
},
);
if (!res.ok) throw new Error(`fixture generate failed: ${res.status}`);
const { data, meta } = await res.json();
console.log(`fixture seed=${meta.seed}`);
return JSON.stringify(data);
}
Replace await readFile("test/fixtures/billing.json", "utf8") in the
Vitest file's beforeAll with await fetchFixture(). The schema still
loads from db/schema.sql. Two things change compared with a committed
file:
- Assertions can't hard-code values when the seed rotates. Compare
the query result with the same rule applied to the fixture in the
test, for example the open invoices filtered from the JSON whose
due_atis before the cut-off, instead of a list of invoice numbers. - Every test file calls the API, and each call counts toward usage.
Fetch once in a
globalSetupand provide the string to the files when many files need the same data.
The run now depends on the network and on the saved template not
changing between the failure and the replay. If the template evolves,
commit the template body and send it to POST /v1/templates/generate
instead, so the version used is in Git next to the test.
Reproducing a CI failure from its seed
With a committed fixture there's nothing to recover: CI and your
machine load identical rows. To reproduce, run the failing class
locally against the same image tag. Pin the tag at least to a major
version, as in postgres:17-alpine, or to a minor version or digest if
you need byte-for-byte the same server. A floating latest tag can
change the database under your tests between two runs.
With fetched data, the seed in the log is the whole recipe. Find the line from the failed job and rerun just that file with it:
FIXTURE_SEED=3141592653 npx vitest run test/invoices.test.ts
The API echoes the seed it used in meta.seed even when the request
didn't send one, which is why the setup logs it on every run. If the
failure reproduces, promote that seed: set it as SEED in
scripts/billing-fixture.sh, point OUT at a second file, and commit
that fixture so the case stays covered. The
deterministic test data post
goes deeper into seed mechanics, including the sources of flakiness a
seed can't remove, like queries without ORDER BY.
Keep fixtures small so containers start fast
Container startup dominates the setup time, not the data. The fixture
in this post is 20 customers and 60 invoices, about 21 KB of JSON,
loaded in two INSERT statements and two setval calls. A few
guidelines keep it that way:
- Size for coverage, not volume. Every state your queries branch on needs a few rows: each status, the empty customer, null dates, boundary dates. Hundreds of rows is plenty for an integration test. Volume and performance belong in a dedicated job, as in test data for load testing.
- One fixture per bounded area. A billing fixture and an inventory fixture load faster and review better than one file for the whole schema. A test class loads only what its queries touch.
- Share containers before you shrink data. If setup is slow, the fix
is usually fewer container starts, through a singleton or
globalSetup, not fewer rows. - Use reusable containers only locally. Testcontainers can keep a
container running between runs with
withReuse. In Java it's experimental, opt-in throughtestcontainers.reuse.enable=truein~/.testcontainers.propertiesor theTESTCONTAINERS_REUSE_ENABLEenvironment variable, and documented as unsuited to CI. A reused database keeps the rows from the last run, so truncate and reload the fixture in setup instead of assuming a fresh one.
CI runners need a Docker environment. GitHub-hosted Ubuntu runners have
one. Where a runner doesn't, Testcontainers can use a remote Docker
host through DOCKER_HOST, or Testcontainers Cloud. If the runner
can't start privileged containers, Ryuk can be turned off with
TESTCONTAINERS_RYUK_DISABLED=true, at the cost of the automatic
cleanup.
FAQ
How do I load test data into a Testcontainers database?
Start the container, create the schema, then insert the data from your
test's setup hook with your normal driver or ORM: @BeforeAll in JUnit
5, beforeAll in Vitest or Jest, or a pytest fixture. Testcontainers
starts and stops the database but doesn't load application data for
you. On PostgreSQL, json_populate_recordset can insert a whole JSON
array into a table in one statement, which keeps the loader to a few
lines.
How do I run an SQL init script with Testcontainers?
In Java, database containers such as PostgreSQLContainer have
withInitScript, which runs a classpath SQL file after the container
starts and before your test gets a connection. In Node, you can copy
the file into the official postgres image's /docker-entrypoint-initdb.d
directory with withCopyFilesToContainer, or simply run the file's
contents with your driver in beforeAll. Init scripts suit schema DDL.
Bulk test data is easier to load from JSON in the setup hook.
Does Testcontainers clean up data between tests?
Not between tests that share a container. Testcontainers removes containers when they are stopped, and its Ryuk sidecar removes leftovers when the test process exits, but rows one test writes stay visible to later tests on the same container. Either give each test class its own container, or run each test in a transaction and roll it back afterwards.
Can Testcontainers tests run in parallel?
Yes, but the Testcontainers JUnit 5 extension is documented as tested
only with sequential execution, so parallelize Java tests with separate
JVM forks, such as Maven Surefire's forkCount or Gradle's
maxParallelForks. Each fork starts its own containers. Vitest and
Jest run test files in separate workers, so a container started in a
file's beforeAll belongs to that file alone and parallel files can't
see each other's rows.
Does JsonFabrica have a Testcontainers module?
No. JsonFabrica is an HTTP API that returns generated JSON. It doesn't
start containers, connect to databases, or write rows. You generate a
fixture with a fixed seed, commit the JSON or fetch it in test setup,
and your own test code inserts it into the container with JDBC, pg,
psycopg, or your ORM.
Testcontainers gives every test run a real database. A seeded template
gives it rows worth testing against: valid foreign keys, planted empty
states, and the same values on every machine. Seeded, parameterized
generation is part of the JsonFabrica API, and the
templates API reference documents the
generate request and its seed and params fields.
Generate realistic test data with JsonFabrica
Describe the shape of your data once, then generate as many fresh, realistic JSON documents as you need via a simple API call.