Add robust local Harbor runtime and CLI interop - #478
Closed
nancyjlau wants to merge 4 commits into
Closed
Conversation
Contributor
|
Preview deployment for your docs. Learn more about Mintlify Previews.
|
nancyjlau
force-pushed
the
nancyjlau/harbor-runner-main
branch
from
July 1, 2026 04:40
6bab8c4 to
f788377
Compare
nancyjlau
force-pushed
the
nancyjlau/harbor-runner-main
branch
from
July 1, 2026 04:50
f788377 to
466cea2
Compare
nancyjlau
force-pushed
the
nancyjlau/harbor-runner-main
branch
from
July 2, 2026 18:34
466cea2 to
4a41ff0
Compare
Replace the Dockerfile-parsing fidelity heuristics (start-script recreation, mkdir-dir restoration, node_modules/vendor/dist submounts, seeded-SQLite restoration, and the hardcoded /app workdir) with a single mechanism: after building the task image, copy its actual working directory onto the host workspace and bind-mount that back over the same guest path. The workspace is then the image's real workdir — source plus every build-generated file, with original mode bits — so nothing the build produced is shadowed by the editable mount, and the guest path is derived from the image's WORKDIR instead of assumed to be /app. This removes ~330 lines of corpus-tuned parsing and makes the runner faithful to any Harbor/terminal-bench image shape rather than the NOVATStyle export specifically. Validated on real Docker across all six artifact classes the heuristics used to cover (compose+postgres, gunicorn start script, node_modules, mkdir dirs, node dist, seeded sqlite): all return reward 0.0 with is_error false and clean teardown. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 27, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
jdchawla29
added a commit
that referenced
this pull request
Jul 28, 2026
integrations/ becomes hud.integrations because an adapted image imports it: the image's CMD serves hud.integrations.harbor:environment, so the module has to ship in the wheel that image installs. Nothing in the SDK or CLI imports it — a Harbor taskset is loaded and placed explicitly. Harbor implements the contract: - load() carries what a task declares, not just its identity: [metadata] as Task.columns (difficulty/category/tags — the platform's facets), [environment] cpu/memory/gpu as RuntimeConfig.resources, real descriptions on templates, verifier.env for the verifier. Time budgets stay off the row — a budget bounds the rollout, not the substrate — and agent_timeout() exposes them for rollout_timeout. - adapt() packages the environment constructor into one HUD-speaking image per build context, whose CMD serves that constructor; rows then run on any container placement. Images are content-addressed and kept, built once per group under a lock. - environment() is what those images serve: the workspace applies the task's declared network isolation, env, workdir and user, healthchecks gate serving, URL MCP servers are published as capabilities, and tests/test.sh grades in place under the task's verifier timeout. Behaviour that cannot be reproduced faithfully is refused, not dropped: allowlist egress, a verifier with its own environment, stdio MCP servers, non-linux os, TPUs, and multi-step tasks (which load, but do not adapt). Environments are grouped by build context *and* declared policy, so one environment always serves one policy. Container, workspace and grading mechanics were ported from #478. Co-authored-by: Nancy <najilau@ucsc.edu>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
HarborRuntimeand routehud eval --format harborthrough it[task]schemaSupported scope
The runtime intentionally fails closed for unsupported Harbor features. The supported subset is Linux, single-step tasks using a
mainservice and shared-containertests/test.shverification. GPU/TPU resources, restricted networking, separate verifier environments, collect hooks, MCP/skills injection, and multi-step/artifact strategies are rejected.Harbor task definitions are trusted input: Dockerfiles, Compose files, and verifier scripts execute against the local Docker daemon. Shared-container verification matches Harbor's shared mode but is not a security boundary against a root agent that leaves background processes behind.
Validation
0, private fixed reward1, authored Compose with healthy sidecar reward1, and official Harbor infra-variable Compose reward1--helptask ID, quoted/multiline shell values, multiline Docker CMD, required JSON build input): reward1.0, zero exceptions778tests pass locally; Ruff, format, Pyright, package build, isolated wheel import, and Harbor schema/publisher validation also passThis remains a draft for maintainer review.