Skip to content

Add robust local Harbor runtime and CLI interop - #478

Closed
nancyjlau wants to merge 4 commits into
mainfrom
nancyjlau/harbor-runner-main
Closed

Add robust local Harbor runtime and CLI interop#478
nancyjlau wants to merge 4 commits into
mainfrom
nancyjlau/harbor-runner-main

Conversation

@nancyjlau

@nancyjlau nancyjlau commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Summary

  • add a real local Docker/Compose-backed HarborRuntime and route hud eval --format harbor through it
  • load current Harbor task metadata and export HUD tasks using Harbor's current [task] schema
  • support Dockerfile, prebuilt-image, and authored-Compose tasks through one Compose lifecycle, including sidecars, users/workdirs, environment interpolation, CPU/memory limits, and agent/verifier timeouts
  • harden Docker execution with bounded output, isolated task environment, secret redaction, process-group cancellation, hidden-test staging, strict reward parsing, and bounded cleanup

Supported scope

The runtime intentionally fails closed for unsupported Harbor features. The supported subset is Linux, single-step tasks using a main service and shared-container tests/test.sh verification. GPU/TPU resources, restricted networking, separate verifier environments, collect hooks, MCP/skills injection, and multi-step/artifact strategies are rejected.

Harbor task definitions are trusted input: Dockerfiles, Compose files, and verifier scripts execute against the local Docker daemon. Shared-container verification matches Harbor's shared mode but is not a security boundary against a root agent that leaves background processes behind.

Validation

  • validated all 620 current Nova Harbor task configs through the runtime loader
  • passed real disposable Docker controls: untouched vulnerable reward 0, private fixed reward 1, authored Compose with healthy sidecar reward 1, and official Harbor infra-variable Compose reward 1
  • passed an installed-Harbor NOP trial on a freshly exported adversarial fixture (custom module/symbol, --help task ID, quoted/multiline shell values, multiline Docker CMD, required JSON build input): reward 1.0, zero exceptions
  • confirmed a startup-timeout control leaves no container or Compose-project leak
  • 778 tests pass locally; Ruff, format, Pyright, package build, isolated wheel import, and Harbor schema/publisher validation also pass
  • GitHub CI passes on Python 3.11 and 3.12, including Ruff, Pyright, and CodeQL

This remains a draft for maintainer review.

@mintlify

mintlify Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated (UTC)
hud 🟢 Ready View Preview Jul 1, 2026, 3:57 AM

@nancyjlau nancyjlau changed the title Refine Harbor runtime CLI integration harbor interop runner addition Jul 1, 2026
nancyjlau and others added 2 commits July 3, 2026 11:15
Replace the Dockerfile-parsing fidelity heuristics (start-script
recreation, mkdir-dir restoration, node_modules/vendor/dist submounts,
seeded-SQLite restoration, and the hardcoded /app workdir) with a single
mechanism: after building the task image, copy its actual working
directory onto the host workspace and bind-mount that back over the same
guest path. The workspace is then the image's real workdir — source plus
every build-generated file, with original mode bits — so nothing the
build produced is shadowed by the editable mount, and the guest path is
derived from the image's WORKDIR instead of assumed to be /app.

This removes ~330 lines of corpus-tuned parsing and makes the runner
faithful to any Harbor/terminal-bench image shape rather than the
NOVATStyle export specifically. Validated on real Docker across all six
artifact classes the heuristics used to cover (compose+postgres,
gunicorn start script, node_modules, mkdir dirs, node dist, seeded
sqlite): all return reward 0.0 with is_error false and clean teardown.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@nancyjlau nancyjlau changed the title harbor interop runner addition Add robust local Harbor runtime and CLI interop Jul 22, 2026
@jdchawla29 jdchawla29 closed this Jul 27, 2026
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29 added a commit that referenced this pull request Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it:
the image's CMD serves hud.integrations.harbor:environment, so the module has
to ship in the wheel that image installs. Nothing in the SDK or CLI imports
it — a Harbor taskset is loaded and placed explicitly.

Harbor implements the contract:

- load() carries what a task declares, not just its identity: [metadata] as
  Task.columns (difficulty/category/tags — the platform's facets),
  [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions
  on templates, verifier.env for the verifier. Time budgets stay off the row
  — a budget bounds the rollout, not the substrate — and agent_timeout()
  exposes them for rollout_timeout.
- adapt() packages the environment constructor into one HUD-speaking image
  per build context, whose CMD serves that constructor; rows then run on any
  container placement. Images are content-addressed and kept, built once per
  group under a lock.
- environment() is what those images serve: the workspace applies the task's
  declared network isolation, env, workdir and user, healthchecks gate
  serving, URL MCP servers are published as capabilities, and tests/test.sh
  grades in place under the task's verifier timeout.

Behaviour that cannot be reproduced faithfully is refused, not dropped:
allowlist egress, a verifier with its own environment, stdio MCP servers,
non-linux os, TPUs, and multi-step tasks (which load, but do not adapt).
Environments are grouped by build context *and* declared policy, so one
environment always serves one policy.

Container, workspace and grading mechanics were ported from #478.

Co-authored-by: Nancy <najilau@ucsc.edu>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants