Testing and validation
Follow the OpenPets quality ladder for desktop tests, package contracts, plugin harnesses, production release gates, and catalog verification.
Testing and validation
OpenPets ships an Electron app, npm packages, third-party-runnable plugin code, and remotely-hosted catalogs - so “does it pass tests” is necessary but not sufficient. This doc lays out the full quality ladder: unit/behavior tests, contract tests at public boundaries, runtime checks, the plugin release validators, and catalog verification - i.e. what “production-valid” means before you ship pets, plugins, packages, or the app.
The quality ladder
From fastest/narrowest to broadest:
- Behavior tests - unit tests of pure logic.
- Contract tests - validate public boundaries (IPC, catalog, manifest) against fixtures so producers and consumers can’t drift apart.
- Runtime checks (
check-*.ts) - assertions about packaging, CSP, SDK conformance, and integration previews that run as part ofcheck/test. - Release validators - the gates that catch production-breaking mistakes
the test suite alone misses (catalog/package drift, missing ZIPs, SHA
mismatches, unresolved
$t:). - Live validation - post-deploy checks against the real origin.
Run the suite with pnpm test (builds first, then each package’s tests) and
pnpm check (per-package typecheck + build + contract checks). See
Development for the command surface.
Desktop tests
The desktop runner (apps/desktop/scripts/run-tests.mjs) orchestrates:
preload syntax checks → test compilation → behavior tests → contract tests →
dist checks. The dist-check step wipes dist/ and rebuilds the full app
(main and renderer), so a finished check leaves a packageable dist/. Three buckets:
- Behavior (
apps/desktop/tests/*.test.ts): lease manager, app state, version checking, ZIP safety, Codex pets, Claude memory, reaction-animation mapping, catalog remote transport/cache behavior and V2/fixture fallback behavior, plugin bridge/gateway guards, andvoice-lifecycle.test.tsfor live microphone-track accounting, capture cancellation/cleanup races, separate timeouts, empty transcripts, and shutdown behavior.remote-control.test.tscovers secure opt-in configuration, verifier-only persistence, authentication, scopes, malformed/oversized requests, rate limiting, rotation, revocation, canonical IPv4/CGNAT boundaries, peer normalization, socket caps/deadlines, away-pet side-effect suppression, and listener shutdown. Compiled to.test-dist/. - Provider-foundation behavior is covered by
provider-profiles.test.ts,provider-presets-and-roles.test.ts,provider-migration.test.ts,provider-credential-deletion.test.ts,provider-service.test.ts, andprovider-configuration-test.test.ts,voice-assistant-host-core.test.ts,voice-realtime-assistant.test.ts,text-model-client.test.ts, andplugin-ai-gateway.test.ts: the canonical adapter/preset catalog, typed profile validation, independent role selection, adapter credential policies, separate text/realtime models, persisted TTS voices and request overrides, versioned migration/defaults/quarantine, URL/header boundaries, fake-endpoint routing, native ElevenLabs Scribe multipart transcription, configured TTS voice previews, temporary draft credentials, native versus compatible codecs, operation snapshots, redacted status, no-fetch unsupported realtime, and provider transport cancellation, bounded-body, malformed-JSON, SSE, and sanitized-error behavior. Tests use fake fetches and no credentials. The provider settings tests also cover atomic save rollback after credential failure, redacted-header add/delete edits, and preservation of secrets outside Control Center snapshots. - Voice Talk behavior tests cover the atomic recording-to-processing transition, double/triple primary toggles without cancellation or parallel turns, capture stop followed by STT/assistant/synthesis, immediate response bubbles before TTS settles, no automatic re-listen after completed speech, explicit activation for the next turn, and the non-destructive native Realtime primary toggle plus its response-completed shutdown boundary. They assert outcomes rather than diagnostic log wording.
- Voice device tests cover preference normalization/persistence, preferred/default/
unavailable resolution, immutable operation snapshots, selected-input propagation
through generic capture and Realtime transport, the truthful unsupported-output
snapshot, and Control Center voice devices UI/preload contracts (
control-center-voice-devices.test.tsandcontrol-center-preload.test.ts). They use injected device lists and do not depend on OS hardware. - Talk/provider lifecycle diagnostics are validated through the existing focused voice and provider behavior tests by preserving their observable cancellation, timeout, output, and cleanup outcomes. Logging tests should not assert exact log wording or raw content; when diagnostics change, validate the relevant focused behavior and the desktop typecheck instead.
- Contract (
apps/desktop/contracts/*.contract.ts): the public boundaries - -catalog-fixture.contract.ts- catalog validation against fixture data.local-ipc-protocol.contract.ts- IPC request/response parsing (IPC and remote control).remote-control-protocol.contract.ts- remote allowlist and secure configuration boundary.plugin-manifest.contract.ts- manifest v1 schema, config refs, permissions, deferred features, action validation (Plugin platform).
- Runtime checks (
apps/desktop/src/check-*.ts): notablycheck-packaging-contract.ts- asserts the packaged app includes every canonical bundled official plugin as extra resources, each plugin’s manifest, entry, declared assets, and locales exist, every unpacked@open-pets/*integration runtime entry (including OpenClaw) exists, and the pet-window CSP allows the bundled emoji font, etc. This is the guard that a packaged build is actually shippable.check-opencode-desktop-setup.ts- verifies the bundled OpenCode setup preview matches expectations.check-zed-desktop.ts- verifies desktop Zed path resolution, preview shape, targeted writes, and removal preservation.
Teams desktop behavior tests
Teams behavior tests cover exact enrollment-link decoding, strict Team Pack payload validation, one-organization state, stable installation metadata, delayed-preview acceptance gating, authoritative identity/expiry refresh, personal/Team ownership isolation, staged reconciliation failure, and Team removal behavior. Enrollment contract tests cover authoritative preview, desktop completion without browser confirmation, same-proof idempotency, incorrect-proof rejection, active device conflicts, and bounded retry. Fake API tests use bounded responses and never persist or assert on bearer credentials.
Package tests & contracts
Each package runs its own check/test. Notable contract/boundary coverage:
packages/client/contracts/client-protocol.contract.ts- the client side of the local IPC and explicit remote-client protocols, paired with the desktop’s server-side contract so both ends validate their separate shapes. Its remote fixture asserts that remote mode works without consulting local discovery.packages/sdk/src/check-plugin-sdk.ts- SDK conformance: compiles/runs a representative plugin against the test harness to detect drift between the published types (index.ts), the harness (testing.ts), and the desktop bridge. Changing the SDK without updating all three fails here. See Plugin SDK v3.packages/openclaw/src/check-openclaw.ts- native OpenClaw package contract: exact management command shapes, supported status/conflict classification, post-install planning, and payload-free non-blocking lifecycle dispatch.@open-pets/dsh- package artifact/load smoke: confirm the built or published artifact loads as the DSH Cordis bundle and its automatic dispatch wiring is available without model tools or MCP setup. See Agent integrations.packages/cursor/src/check-cursor.ts,packages/zed/src/check-zed.ts,packages/opencodechecks, etc. - validate the safe config-write behavior (status classification, redaction, symlink/oversize rejection, atomic writes, interrupted-write recovery, uninstall preserving user entries). See Agent integrations.
Plugin testing
-
Unit: each official plugin has a
test.jsusing@open-pets/plugin-sdk/testing- fake time/events, descriptor-level assertions, no Electron. Run viapnpm plugins:test, which first runspnpm plugins:locales(scripts/check-plugin-locales.mjs) to verify every$t:/ctx.t()key resolves. See Plugin SDK v3. -
Manifest validation:
openpets plugin validate <dir>checks manifest, permissions, SDK compatibility, config field types, network hosts, asset formats/size caps, entry files, and panels - run it before packaging. -
Calendar Airmail: its deterministic harness coverage should exercise the primary-calendar reconciliation, state-appropriate connection commands, ten-minute and start deliveries, duplicate suppression, selected/default couriers, and reconnect-required cleanup. Run its plugin test alongside
pnpm plugins:locales,pnpm plugins:test, andpnpm --filter @open-pets/plugin-sdk checkwhen changing its SDK-facing behavior. -
Delivery/picker boundary: desktop bridge tests cover
ui:deliverypermission and lifecycle semantics; manifest validation covers declared sprite-grid options and asset references. For an Electron end-to-end smoke run, verify that the Airmail settings grid loads each bundled courier, keyboard and pointer selection persist, reduced motion is static, and a test delivery uses the selected courier without requiring any installed pet.
Plugin release validation (production gate)
plugins:check alone is not release-readiness. The dedicated validators are
the production gate (scripts/validate-plugin-release.mjs):
| Command | When | Catches |
|---|---|---|
pnpm plugins:package |
build artifacts | (produces catalog + ZIP staging) |
pnpm plugins:validate-release |
before deploy | unresolved $t: names/descriptions in catalog cards, missing plugin ZIPs, SHA mismatches, missing locales/en.json, missing declared assets/entry files (including courier sprites), catalog/package drift, and community plugin sidecar validation (provenance.json, submissions.json) |
pnpm plugins:validate-live |
after deploy/R2 upload | the same, against the live catalog + live ZIPs & live sidecars |
Plugin sidecar validation
The release validator automatically loads web/public/plugins/provenance.json
and web/public/plugins/submissions.json and asserts:
- Every community plugin mapped in the catalog has a matching provenance entry.
- All provenance entries contain valid URLs, hex SHAs (40 characters), and formatted dates.
- Update policy is strictly limited to either
safe-autoormanual-review. - Pending submissions are well-formed and are not also present in the installable catalog.
The full pre-ship sequence (from AGENTS.md):
pnpm plugins:package → pnpm plugins:validate-release → deploy/upload →
pnpm plugins:validate-live. Treat a failing validator as a hard stop - these
are exactly the mistakes that 404 a plugin or render a raw $t:... to users.
For the full plugin catalog release path in one command, run
pnpm plugins:release; it packages, validates, publishes ZIPs, deploys the web
catalog, then validates the live catalog.
Catalog verification (production gate for pets)
Pet catalogs have a parallel “doctor” run from web/ (read-only; safe anytime).
The gate in brief:
| Command | Adds |
|---|---|
bun run verify:catalog |
manifest integrity, artifact freshness vs manifest, on-disk assets, orphan dirs |
bun run verify:catalog:remote |
range-validates every ZIP referenced by the local v3 catalog against the desktop install contract |
bun run verify:catalog:prod |
fetches deployed v3 pages, validates every deployed ZIP, and diffs local vs prod (pending/removed) |
bun run verify:catalog:all |
all local, local-catalog remote, deployed-prod ZIP, and drift checks |
The non-negotiable rule: never ship a catalog entry whose ZIP is missing or
fails the desktop install contract. Run verify:catalog:remote before
deploying and verify:catalog:prod after to confirm both the deployed catalog
and its ZIPs are valid. The ZIP checks use bounded HTTP range reads of the
central directory; they do not download spritesheet payloads. See
Catalogs.
What “production-valid” means
Before shipping, the relevant gate must be green:
- A package change →
pnpm check+pnpm test(incl. contract + conformance checks) pass; for@open-pets/dsh, the package artifact/load smoke also passes. - An app change → desktop behavior + contract + runtime checks pass; if it
touches packaging/CSP/bundled plugins,
check-packaging-contract.tspasses. A delivery, trusted-asset protocol, or sprite-picker change additionally needs the desktop bridge/static checks and the targeted Electron smoke above. - A plugin release →
validate-releasebefore deploy,validate-liveafter. - A pet catalog change →
verify:catalog:remotebefore,verify:catalog:prodafter; both validate the relevant v3 ZIP archive contracts. - Linux-specific behavior → validated on the Ubuntu VM (Development).
The local desktop release pipeline runs the packaged-output contract against a
fresh unpacked build for every platform/architecture target and against the
extracted payload of the actual distributable before it reaches copy, tag, or
publication stages. Externally staged DEB/RPM payloads are extracted and checked
before copy. Artifact name, size, and checksum checks do not replace this
validation; target selection is explicit so Linux packages require the matching
sharp-linux-* native runtime rather than the release host’s runtime. The
macOS release host uses unsquashfs (install with brew install squashfs) and
the ELF section table to extract the actual Linux Type-2 AppImage payload without
executing its Linux binary. RPM payloads are first spooled to a temporary CPIO
file, avoiding a release-host memory limit on full desktop packages. The
focused packaging-contract.test.ts fixture regression proves that staged DEB
and RPM payloads missing openpets.system-resources or the target Sharp runtime
are hard failures. Release checkpoints also persist SHA-256 digests for every
artifact output; release-checkpoint.test.mjs proves that a same-size mutation
stales the stage while an unchanged checkpoint remains reusable.
Release gates and ordering
Release commands do not replace the quality gates:
- Package gate →
pnpm checkandpnpm testpass for the workspace and the current public package plan. - Desktop gate →
pnpm --filter @open-pets/desktop checkandpnpm --filter @open-pets/desktop testpass, with the workspace build required by the desktop release flow.
For a full shared-version release, publish the complete dynamic package plan
with pnpm release:npm -- --yes, verify every planned package/version on the
public npm registry, and only then promote the desktop tag. The npm helper’s
printed plan is authoritative; this document must not become a stale package
inventory. A partial npm publish is recoverable by rerunning the same command;
already published versions are skipped.
A desktop-only release uses the Desktop gate and does not publish npm packages,
unless it introduces a new exact npm integration version. In that case, publish
and verify that version before the desktop tag is promoted. Never use
--skip-checks on a live npm or desktop release. For historical npm recovery,
use the tagged source explicitly:
pnpm release:npm -- --yes --ref vX.Y.Z
If a gate is skipped, say so explicitly rather than implying coverage. Contract and validator failures are signal, not noise - they encode the ways this product has broken in production before.
