Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
115 commits
Select commit Hold shift + click to select a range
e268794
feat: add stock kimi tool-use eval
adibarra Aug 10, 2026
b7d0d7e
refactor: simplify and expose tool-use eval
adibarra Aug 10, 2026
840537b
chore: merge latest main for release
adibarra Aug 10, 2026
da29994
refactor: isolate and clarify verifier integration
adibarra Aug 10, 2026
be0ca08
chore: merge latest main before release
adibarra Aug 10, 2026
b0fd8cc
fix: distinguish successful failure artifact writes
adibarra Aug 10, 2026
1684f55
fix: preserve agentic eval decoding mode
adibarra Aug 10, 2026
0223399
chore: merge latest main before release
adibarra Aug 10, 2026
5885859
fix: complete eval-only workflow runs
adibarra Aug 10, 2026
83d7f57
revert: keep eval-only collection scoped
adibarra Aug 10, 2026
ee80bce
chore: merge latest main before review
adibarra Aug 10, 2026
07dc9d9
fix: correct verifier failure metadata and links
adibarra Aug 10, 2026
6ec09bb
chore: preserve existing collector test formatting
adibarra Aug 10, 2026
2b6e5e8
fix: preserve intended verifier sample count
adibarra Aug 10, 2026
5e3bc30
chore: merge latest main before final review
adibarra Aug 11, 2026
6793c9e
fix: preserve configurable eval dispatch behavior
adibarra Aug 11, 2026
4ed2975
Merge remote-tracking branch 'origin/main' into feat/tool-use-eval-smoke
adibarra Aug 11, 2026
294e39d
fix: harden verifier review paths
adibarra Aug 11, 2026
ce37f46
fix: scope eval state and links
adibarra Aug 11, 2026
405dc0e
ci: exclude faulty b300 node from slurm
adibarra Aug 11, 2026
847982f
ci: apply B300 node exclusions to allocations
adibarra Aug 11, 2026
15b0857
feat: enable multinode kimi verifier
adibarra Aug 12, 2026
13c5a45
fix: harden Kimi eval runtime failures
adibarra Aug 12, 2026
134906e
test: capture Kimi response diagnostics
adibarra Aug 12, 2026
23851fb
Revert "test: capture Kimi response diagnostics"
adibarra Aug 12, 2026
9a661a6
test: capture deterministic Kimi diagnostics
adibarra Aug 12, 2026
53b3ca8
fix: enable Kimi structural tool constraints
adibarra Aug 12, 2026
6d08f1e
chore: format Kimi recipe regression
adibarra Aug 12, 2026
f1fb29d
fix: retry transient Kimi verifier downloads
adibarra Aug 12, 2026
0e679c5
fix: stabilize and clean Kimi verifier
adibarra Aug 12, 2026
34986b9
Merge remote-tracking branch 'origin/main' into feat/tool-use-eval-smoke
adibarra Aug 12, 2026
37e7363
Merge remote-tracking branch 'origin/main' into feat/tool-use-eval-smoke
adibarra Aug 14, 2026
8952cdc
fix: preserve tool eval dispatch contracts
adibarra Aug 14, 2026
ba5b363
fix: enable structured Qwen tool calls
adibarra Aug 14, 2026
a4f69b9
fix: preserve single-node eval topology metadata
adibarra Aug 14, 2026
333a032
test: cover isolated verifier runtime bootstrap
adibarra Aug 14, 2026
26c0154
fix: configure DSV4 tool call parsers
adibarra Aug 14, 2026
a2ad629
fix: configure sglang tool call parsers
adibarra Aug 14, 2026
ffe9054
fix: harden tool eval dispatch and artifacts
adibarra Aug 14, 2026
209fbf4
fix: preserve selected SRT eval concurrency
adibarra Aug 14, 2026
38a867e
fix: preserve dsv4 thinking request mode
adibarra Aug 14, 2026
6d19146
chore: merge main into verifier branch
adibarra Aug 14, 2026
ad9512b
feat: add minimax m3 provider smoke
adibarra Aug 14, 2026
301dcdf
fix: preserve minimax m3 generation budget
adibarra Aug 14, 2026
165135e
fix: retry truncated minimax response bodies
adibarra Aug 14, 2026
ef08f9b
fix: narrow minimax smoke to schema case
adibarra Aug 15, 2026
f6cb909
fix: remove obsolete gb200 recipe field
adibarra Aug 15, 2026
439d884
fix: enable frontend tool call parsing
adibarra Aug 15, 2026
4fed421
fix: parse GB200 vendor tool calls
adibarra Aug 15, 2026
6f62744
fix: force guided Kimi tool calls
adibarra Aug 15, 2026
fbb8fe2
fix: guide all GB200 Kimi recipes
adibarra Aug 15, 2026
92c8c4f
fix: parse GB200 Kimi tool calls
adibarra Aug 15, 2026
49fd66f
chore: merge current main
adibarra Aug 15, 2026
7edb59b
fix: harden vendor artifact cleanup
adibarra Aug 15, 2026
ade0abd
fix: harden transient endpoint registration retries
adibarra Aug 15, 2026
2567569
fix: allow trusted Kimi frontend code
adibarra Aug 15, 2026
a847479
feat: add deterministic bfcl tool smoke
adibarra Aug 15, 2026
9fff29e
fix: install bfcl audio dependency
adibarra Aug 15, 2026
2550733
fix: require externally grounded bfcl calls
adibarra Aug 15, 2026
2dec376
fix: harden tool eval failure handling
adibarra Aug 15, 2026
592215f
feat: add opt-in full verifier suites
adibarra Aug 15, 2026
830c05a
feat: add recommended BFCL eval suites
adibarra Aug 16, 2026
72a8a6a
docs: document BFCL model quality suites
adibarra Aug 16, 2026
61e84e4
docs: clarify BFCL suite coverage limits
adibarra Aug 16, 2026
1a8cfdc
feat: preserve BFCL license with artifacts
adibarra Aug 16, 2026
a4713af
chore: merge main into tool-use experiment
adibarra Aug 17, 2026
7b61603
feat: complete tool-use eval integration
adibarra Aug 17, 2026
f020a8e
fix: enable tool parsing for kimi k3
adibarra Aug 17, 2026
b57d8bb
Merge branch 'main' into experiment/tool-use-eval-full
adibarra Aug 18, 2026
d070e3b
fix: harden tool use deployment verification
adibarra Aug 18, 2026
b8d681f
fix: normalize infrastructure failure sample counts
adibarra Aug 18, 2026
b816503
ci: repair amd workspace ownership
adibarra Aug 18, 2026
6cc9937
fix: preserve kimi quality failures
adibarra Aug 18, 2026
240a564
fix: harden bfcl transport retries
adibarra Aug 18, 2026
c63d0d2
chore: merge current main branch
adibarra Aug 28, 2026
708c8b5
fix: preserve stock tool use validators
adibarra Aug 29, 2026
80fde1c
feat: select exact generated deployments for eval shards
adibarra Aug 29, 2026
07bfc0d
chore: remove unintended generator formatting churn
adibarra Aug 29, 2026
69e9e1d
fix: harden cached image and recipes
adibarra Aug 29, 2026
c813727
fix: restore GB200 draft model source
adibarra Aug 29, 2026
aac233a
fix: gate tool evaluations on readiness
adibarra Aug 29, 2026
b74fe20
fix: require registered model before evaluation
adibarra Aug 29, 2026
533a027
fix: wait for active chat route
adibarra Aug 29, 2026
8ef8107
fix: gate eval dispatcher on route
adibarra Aug 30, 2026
0bfb0fb
fix: preserve selected TRT eval framework
adibarra Aug 30, 2026
b6d4085
fix: send model-aware readiness probe
adibarra Aug 30, 2026
529b343
fix: probe chat route registration
adibarra Aug 30, 2026
bc097ac
fix: stabilize registered models before eval
adibarra Aug 30, 2026
089e707
fix: accept stable OpenAI server readiness
adibarra Aug 30, 2026
c12cdd6
fix: stabilize OpenAI health before eval
adibarra Aug 30, 2026
4961aec
fix: refresh expired MiniMax image
adibarra Aug 30, 2026
f76459e
fix: allocate hybrid DP ranks correctly
adibarra Aug 30, 2026
dcc2763
fix: accept stock TensorRT chat requests
adibarra Aug 30, 2026
5aec415
fix: keep AgentX dependency setup rootless
adibarra Aug 30, 2026
b4fdc49
fix: support packaged MiniMax evaluator imports
adibarra Aug 30, 2026
19bb70a
fix: size vLLM offload allocations correctly
adibarra Aug 30, 2026
64e85a8
fix: split heterogeneous vLLM cache layers
adibarra Aug 30, 2026
6b0c5db
chore: normalize vLLM offload patcher imports
adibarra Aug 30, 2026
bd13916
fix: patch B300 vLLM offload startup
adibarra Aug 30, 2026
c75ca23
chore: merge main into tool evals
adibarra Aug 31, 2026
d5bde4c
fix: update validators and repair smoke
adibarra Aug 31, 2026
38dd382
refactor: narrow tool parser evaluation scope
adibarra Aug 31, 2026
1c0d6f2
docs: document runtime patch removal gates
adibarra Aug 31, 2026
dc3cd43
chore: merge latest main into evaluations
adibarra Aug 31, 2026
03caa00
docs: link upstream runtime patch tracking
adibarra Aug 31, 2026
139128e
fix: cut MI300X routing over to the new AMD cluster
cquil11 Aug 31, 2026
260c5b9
fix: classify empty tool outputs correctly
adibarra Aug 31, 2026
8b7037d
chore: merge latest main before validation
adibarra Aug 31, 2026
2249e46
fix: bound full BFCL evaluation runtime
adibarra Sep 1, 2026
c4cb22c
ci: make tool evaluation scores nonblocking
adibarra Sep 1, 2026
3c2a9e9
feat: auto-run matching model vendor validators
adibarra Sep 1, 2026
f967d65
fix: avoid duplicate eval matrix scoring
adibarra Sep 1, 2026
3f8eaf9
fix: preserve agentic eval topology settings
adibarra Sep 1, 2026
f672621
fix: preserve GB200 synthetic acceptance dispatch
adibarra Sep 1, 2026
1cf10bf
docs: remove unnecessary PR waiver files
adibarra Sep 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 22 additions & 1 deletion .github/workflows/benchmark-multinode-tmpl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -142,6 +142,16 @@ on:
type: boolean
required: false
default: false
eval-framework:
description: "Eval runner (lm-eval, swebench, kimi-vendor, minimax-vendor, or bfcl)"
type: string
required: false
default: "lm-eval"
eval-suite:
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
type: string
required: false
default: ""
eval-conc:
description: "Concurrency value or space-separated list for eval requests (overrides default max-of-conc-list)"
type: string
Expand Down Expand Up @@ -242,6 +252,8 @@ env:
DECODE_HARDWARE: ${{ inputs.decode-hardware }}
RUN_EVAL: ${{ inputs.run-eval }}
EVAL_ONLY: ${{ inputs.eval-only }}
EVAL_FRAMEWORK: ${{ inputs.eval-framework }}
EVAL_SUITE: ${{ inputs.eval-suite }}
EVAL_CONC: ${{ inputs.eval-conc }}
EVAL_LIMIT: ${{ inputs.eval-limit }}
SWEBENCH_GEN_MODE: ${{ inputs.swebench-gen-mode }}
Expand Down Expand Up @@ -403,6 +415,9 @@ jobs:
fi
# Export RESULT_FILENAME early so it's available for artifact uploads even if cancelled
echo "RESULT_FILENAME=${RESULT_FILENAME}" >> "$GITHUB_ENV"
echo "EVAL_ARTIFACT_RECIPE=${RECIPE_FINGERPRINT:0:16}" >> "$GITHUB_ENV"
eval_artifact_conc="$(python3 -c 'import hashlib,os; print(hashlib.sha256(os.environ["CONC_LIST"].encode()).hexdigest()[:12])')"
echo "EVAL_ARTIFACT_CONC=${eval_artifact_conc}" >> "$GITHUB_ENV"

export ${{ join(fromJson(inputs.prefill-additional-settings), ' ') }} ${{ join(fromJson(inputs.decode-additional-settings), ' ') }}
export IS_MULTINODE=true
Expand Down Expand Up @@ -528,14 +543,17 @@ jobs:
if: ${{ always() && (env.RUN_EVAL == 'true' || inputs.eval-only) }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: eval_${{ env.RESULT_FILENAME }}
name: eval_${{ env.EXP_NAME }}_${{ env.PRECISION }}_${{ env.FRAMEWORK }}_p${{ env.PREFILL_NUM_WORKERS }}x${{ env.PREFILL_TP }}p${{ env.PREFILL_PP_SIZE }}c${{ env.PREFILL_DCP_SIZE }}k${{ env.PREFILL_PCP_SIZE }}e${{ env.PREFILL_EP }}d${{ env.PREFILL_DP_ATTN }}_d${{ env.DECODE_NUM_WORKERS }}x${{ env.DECODE_TP }}p${{ env.DECODE_PP_SIZE }}c${{ env.DECODE_DCP_SIZE }}k${{ env.DECODE_PCP_SIZE }}e${{ env.DECODE_EP }}d${{ env.DECODE_DP_ATTN }}_g${{ env.DISAGG }}_r${{ env.EVAL_ARTIFACT_RECIPE }}_c${{ env.EVAL_ARTIFACT_CONC }}_kv${{ env.KV_OFFLOADING }}-${{ env.KV_OFFLOAD_BACKEND }}_spec${{ env.SPEC_DECODING }}_${{ env.EVAL_FRAMEWORK }}_${{ env.EVAL_SUITE }}_${{ runner.name }}_${{ github.run_attempt }}
path: |
meta_env.json
results*.json
*_report.json
*_results.jsonl
sample*.jsonl
agent_preds.json
predictions.jsonl
swebench_report_*.json
*_artifacts.tar.gz
*.traj*
if-no-files-found: ${{ inputs.eval-only && 'error' || 'ignore' }}

Expand All @@ -553,6 +571,9 @@ jobs:
run: |
rm -f meta_env.json || true
rm -f results*.json || true
rm -f -- ./*_report.json || true
rm -f -- ./*_results.jsonl || true
rm -f -- ./*_artifacts.tar.gz || true
rm -f sample*.jsonl || true
rm -f agent_preds.json predictions.jsonl swebench_report_*.json *.traj* || true

Expand Down
29 changes: 27 additions & 2 deletions .github/workflows/benchmark-tmpl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -90,6 +90,16 @@ on:
type: boolean
required: false
default: false
eval-framework:
description: "Eval runner (lm-eval, swebench, kimi-vendor, minimax-vendor, or bfcl)"
type: string
required: false
default: "lm-eval"
eval-suite:
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
type: string
required: false
default: ""
random-range-ratio:
required: false
type: string
Expand Down Expand Up @@ -179,6 +189,8 @@ env:
DISAGG: ${{ inputs.disagg }}
RUN_EVAL: ${{ inputs.run-eval }}
EVAL_ONLY: ${{ inputs.eval-only }}
EVAL_FRAMEWORK: ${{ inputs.eval-framework }}
EVAL_SUITE: ${{ inputs.eval-suite }}
# Agentic-coding env. Fixed-seq-len jobs leave these empty.
SCENARIO_TYPE: ${{ inputs.scenario-type }}
SCENARIO_SUBDIR: ${{ inputs.scenario-type == 'agentic-coding' && 'agentic/' || 'fixed_seq_len/' }}
Expand All @@ -203,7 +215,7 @@ env:
MODAL_TOKEN_ID: ${{ secrets.MODAL_TOKEN_ID }}
MODAL_TOKEN_SECRET: ${{ secrets.MODAL_TOKEN_SECRET }}
# These b300 nodes are currently broken.
SALLOC_EXCLUDE: 'b300-005,b300-006'
SALLOC_EXCLUDE: 'b300-005,b300-006,b300-017'

permissions:
contents: read
Expand Down Expand Up @@ -263,6 +275,13 @@ jobs:
steps:
- name: Resource cleanup (pre-run)
run: &resource-cleanup |
# Containerized AMD Slurm jobs can leave root-owned artifacts when
# interrupted before their launcher EXIT trap runs. Repair ownership
# before checkout cleanup and again after the job via this shared step.
if [[ "${{ inputs.runner }}" == "cluster:mi355x-amds" && -d "$GITHUB_WORKSPACE" ]]; then
sudo chown -R "$(id -u):$(id -g)" "$GITHUB_WORKSPACE"
fi

# Cleanup Docker resources
if command -v docker >/dev/null 2>&1 && docker info >/dev/null 2>&1; then
echo "[Docker] Cleaning up resources ..."
Expand Down Expand Up @@ -443,10 +462,13 @@ jobs:
if: ${{ always() && (env.RUN_EVAL == 'true' || inputs.eval-only) }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: eval_${{ env.EXP_NAME }}_${{ env.RESULT_FILENAME }}
name: eval_${{ env.RESULT_FILENAME }}_${{ env.EVAL_FRAMEWORK }}_${{ env.EVAL_SUITE }}_${{ github.run_attempt }}
path: |
meta_env.json
results*.json
*_report.json
*_results.jsonl
*_artifacts.tar.gz
sample*.jsonl
agent_preds.json
predictions.jsonl
Expand All @@ -464,7 +486,10 @@ jobs:
rm -f meta_env.json || true
# Remove any eval results JSONs that were moved into workspace
rm -f results*.json || true
rm -f -- ./*_report.json || true
rm -f -- ./*_results.jsonl || true
rm -f sample*.jsonl || true
rm -f -- ./*_artifacts.tar.gz || true
rm -f agent_preds.json predictions.jsonl swebench_report_*.json *.traj* || true

- name: Resource cleanup (post-run)
Expand Down
59 changes: 53 additions & 6 deletions .github/workflows/e2e-tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,16 @@ on:
required: false
type: string
default: ""
eval-framework:
description: "Eval runner override (auto uses model-aware selection; lm-eval, swebench, kimi-vendor, minimax-vendor, or bfcl)"
required: false
type: string
default: "auto"
eval-suite:
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
required: false
type: string
default: ""
swebench-gen-mode:
description: "SWE-bench generation mode (single-shot | agentic). Empty = agentic (single-shot is an explicit debugging escape hatch)."
required: false
Expand Down Expand Up @@ -125,6 +135,16 @@ on:
required: false
type: string
default: ""
eval-framework:
description: "Eval runner override (auto uses model-aware selection; lm-eval, swebench, kimi-vendor, minimax-vendor, or bfcl)"
required: false
type: string
default: "auto"
eval-suite:
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
required: false
type: string
default: ""
swebench-gen-mode:
description: "SWE-bench generation mode (single-shot | agentic). Empty = agentic (single-shot is an explicit debugging escape hatch)."
required: false
Expand Down Expand Up @@ -236,16 +256,26 @@ jobs:
CMD+=(--evals-only)
fi
RAW_CONFIG_JSON=$("${CMD[@]}")
CONFIG_JSON=$(python3 -c 'import json,sys; data=json.load(sys.stdin); rows=[row for family in ("single_node","multi_node") for group in data.get(family,{}).values() for row in group]; rows.extend(row for family in ("evals","agentic_evals","multinode_evals") for row in data.get(family,[])); print(json.dumps(rows))' <<<"$RAW_CONFIG_JSON")
CONFIG_JSON=$(python3 -c 'import json,sys; data=json.load(sys.stdin); rows=[row for family in ("single_node","multi_node") for group in data.get(family,{}).values() for row in group]; rows.extend(row for family in ("evals","agentic_evals","multinode_evals","multinode_agentic_evals") for row in data.get(family,[])); print(json.dumps(rows))' <<<"$RAW_CONFIG_JSON")
else
GENERATE_COMMAND="${{ inputs.generate-cli-command || github.event.inputs.generate-cli-command }}"
if [ -z "$GENERATE_COMMAND" ]; then
echo "generate-cli-command is required outside trusted changelog dispatch mode" >&2
exit 1
fi
read -r -a GENERATE_ARGS <<< "$GENERATE_COMMAND"
if [ "$TRIM_CONC" = "true" ]; then
GENERATE_ARGS+=(--trim-conc)
fi
if [ "$ALL_EVALS" = "true" ]; then
GENERATE_ARGS+=(--all-evals)
fi
if [ "$EVALS_ONLY" = "true" ]; then
GENERATE_ARGS+=(--evals-only)
fi
CONFIG_JSON=$(uv run --no-project --with pydantic --with pyyaml --python 3.12 \
"${GITHUB_WORKSPACE}/utils/matrix_logic/generate_sweep_configs.py" \
$GENERATE_COMMAND)
"${GENERATE_ARGS[@]}")
fi
score_matrix() {
local family="$1"
Expand All @@ -255,9 +285,9 @@ jobs:
--queue-namespace "${{ github.run_id }}:${{ github.run_attempt }}:${family}" \
--labels-json "$PR_LABELS"
}
AGENTIC=$(echo "$CONFIG_JSON" | python3 -c "import sys,json; d=json.load(sys.stdin); print(json.dumps([x for x in d if x.get('scenario-type') == 'agentic-coding' and 'prefill' not in x and not x.get('run-eval', False)]))" | score_matrix agentic)
AGENTIC=$(echo "$CONFIG_JSON" | python3 -c "import sys,json; d=json.load(sys.stdin); print(json.dumps([x for x in d if x.get('scenario-type') == 'agentic-coding' and 'prefill' not in x and not x.get('eval-only', False)]))" | score_matrix agentic)
AGENTIC_EVAL=$(echo "$CONFIG_JSON" | python3 -c "import sys,json; d=json.load(sys.stdin); print(json.dumps([x for x in d if x.get('scenario-type') == 'agentic-coding' and 'prefill' not in x and x.get('run-eval', False)]))" | score_matrix agentic-eval)
MULTI_AGENTIC=$(echo "$CONFIG_JSON" | python3 -c "import sys,json; d=json.load(sys.stdin); print(json.dumps([x for x in d if x.get('scenario-type') == 'agentic-coding' and 'prefill' in x and not x.get('run-eval', False)]))" | score_matrix multi-agentic)
MULTI_AGENTIC=$(echo "$CONFIG_JSON" | python3 -c "import sys,json; d=json.load(sys.stdin); print(json.dumps([x for x in d if x.get('scenario-type') == 'agentic-coding' and 'prefill' in x and not x.get('eval-only', False)]))" | score_matrix multi-agentic)
MULTI_AGENTIC_EVAL=$(echo "$CONFIG_JSON" | python3 -c "import sys,json; d=json.load(sys.stdin); print(json.dumps([x for x in d if x.get('scenario-type') == 'agentic-coding' and 'prefill' in x and x.get('run-eval', False)]))" | score_matrix multi-agentic-eval)
SINGLE=$(echo "$CONFIG_JSON" | python3 -c "import sys,json; d=json.load(sys.stdin); print(json.dumps([x for x in d if 'prefill' not in x and x.get('scenario-type') != 'agentic-coding' and not x.get('eval-only', False)]))" | score_matrix single)
MULTI=$(echo "$CONFIG_JSON" | python3 -c "import sys,json; d=json.load(sys.stdin); print(json.dumps([x for x in d if 'prefill' in x and x.get('scenario-type') != 'agentic-coding' and not x.get('eval-only', False)]))" | score_matrix multi)
Expand Down Expand Up @@ -381,7 +411,11 @@ jobs:
decode-additional-settings: ${{ toJson(matrix.config.decode.additional-settings) }}
run-eval: true
eval-only: true
eval-conc: ${{ matrix.config['eval-all-concs'] && join(matrix.config.conc, ' ') || matrix.config['eval-conc'] }}
eval-conc: ${{ (inputs.eval-framework == 'auto' && (matrix.config['eval-framework'] || 'lm-eval') || inputs.eval-framework) == 'lm-eval' && matrix.config['eval-all-concs'] && join(matrix.config.conc, ' ') || matrix.config['eval-conc'] }}
eval-limit: ${{ inputs.eval-limit }}
eval-framework: ${{ inputs.eval-framework == 'auto' && (matrix.config['eval-framework'] || 'lm-eval') || inputs.eval-framework }}
eval-suite: ${{ inputs.eval-suite != '' && inputs.eval-suite || (inputs.eval-framework == 'auto' && matrix.config['eval-suite'] || '') }}
swebench-gen-mode: ${{ inputs.swebench-gen-mode }}
ref: ${{ inputs.ref }}

test-sweep-agentic:
Expand Down Expand Up @@ -451,12 +485,18 @@ jobs:
model-prefix: ${{ matrix.config.model-prefix }}
framework: ${{ matrix.config.framework }}
precision: ${{ matrix.config.precision }}
router: ${{ matrix.config.router && toJson(matrix.config.router) || '' }}
kv-p2p-transfer: ${{ matrix.config['kv-p2p-transfer'] || '' }}
tp: ${{ matrix.config.tp }}
pp: ${{ matrix.config.pp }}
dcp-size: ${{ matrix.config.dcp-size }}
pcp-size: ${{ matrix.config.pcp-size }}
ep: ${{ matrix.config.ep }}
dp-attn: ${{ matrix.config.dp-attn }}
conc: ${{ matrix.config.conc }}
kv-offloading: ${{ matrix.config.kv-offloading }}
kv-offload-backend: ${{ matrix.config['kv-offload-backend'].name }}
kv-offload-backend-metadata: ${{ matrix.config['kv-offload-backend'] && toJson(matrix.config['kv-offload-backend']) || '' }}
total-cpu-dram-gb: ${{ matrix.config.total-cpu-dram-gb }}
duration: ${{ inputs.agentx-fast && '1200' || (inputs.duration-override != '' && inputs.duration-override || matrix.config.duration) }}
agentx-fast: ${{ inputs.agentx-fast }}
Expand All @@ -469,6 +509,8 @@ jobs:
eval-only: true
eval-limit: ${{ inputs.eval-limit }}
swebench-gen-mode: ${{ inputs.swebench-gen-mode }}
eval-framework: ${{ inputs.eval-framework == 'auto' && (matrix.config['eval-framework'] || 'lm-eval') || inputs.eval-framework }}
eval-suite: ${{ inputs.eval-suite != '' && inputs.eval-suite || (inputs.eval-framework == 'auto' && matrix.config['eval-suite'] || '') }}
scenario-type: agentic-coding
ref: ${{ inputs.ref }}

Expand Down Expand Up @@ -559,7 +601,7 @@ jobs:
precision: ${{ matrix.config.precision }}
router: ${{ matrix.config.router && toJson(matrix.config.router) || '' }}
kv-p2p-transfer: ${{ matrix.config['kv-p2p-transfer'] || '' }}
conc-list: ${{ toJson(matrix.config.conc) }}
conc-list: ${{ format('[{0}]', matrix.config['eval-conc']) }}
spec-decoding: ${{ matrix.config.spec-decoding }}
disagg: ${{ matrix.config.disagg }}
prefill-hardware: ${{ matrix.config.prefill.hardware }}
Expand Down Expand Up @@ -590,6 +632,8 @@ jobs:
eval-conc: ${{ matrix.config['eval-conc'] }}
eval-limit: ${{ inputs.eval-limit }}
swebench-gen-mode: ${{ inputs.swebench-gen-mode }}
eval-framework: ${{ inputs.eval-framework == 'auto' && (matrix.config['eval-framework'] || 'lm-eval') || inputs.eval-framework }}
eval-suite: ${{ inputs.eval-suite != '' && inputs.eval-suite || (inputs.eval-framework == 'auto' && matrix.config['eval-suite'] || '') }}
scenario-type: agentic-coding
ref: ${{ inputs.ref }}

Expand Down Expand Up @@ -670,6 +714,9 @@ jobs:
run-eval: true
eval-only: true
eval-limit: ${{ inputs.eval-limit }}
eval-framework: ${{ inputs.eval-framework == 'auto' && (matrix.config['eval-framework'] || 'lm-eval') || inputs.eval-framework }}
eval-suite: ${{ inputs.eval-suite != '' && inputs.eval-suite || (inputs.eval-framework == 'auto' && matrix.config['eval-suite'] || '') }}
swebench-gen-mode: ${{ inputs.swebench-gen-mode }}
ref: ${{ inputs.ref }}

collect-results:
Expand Down
14 changes: 7 additions & 7 deletions .github/workflows/run-sweep.yml
Original file line number Diff line number Diff line change
Expand Up @@ -864,12 +864,12 @@ jobs:
eval-only: true
eval-conc: ${{ matrix.config['eval-all-concs'] && join(matrix.config.conc, ' ') || matrix.config['eval-conc'] }}

# Multi-node agentic (SWE-bench) eval rows carry the agentic input shape,
# so they are dispatched with sweep-multi-node-agentic's inputs rather
# than sweep-multi-node-evals' fixed-seq-len inputs (isl/osl/max-model-len,
# which agentic rows don't have). SWE-bench doesn't support batched
# concurrencies (unlike lm-eval), so eval-conc is always a single value,
# never the joined-list form sweep-multi-node-evals uses.
# Multi-node agentic GSM8K eval rows carry the agentic input shape, so
# they are dispatched with sweep-multi-node-agentic's inputs rather than
# sweep-multi-node-evals' fixed-seq-len inputs (isl/osl/max-model-len,
# which agentic rows don't have). Agentic selection uses one highest
# eval-conc per topology; fixed-sequence --all-evals rows may instead pass
# a joined concurrency list to lm-eval.
sweep-multi-node-agentic-evals:
needs: [setup, canary-select, canary-sweep]
if: >-
Expand Down Expand Up @@ -906,7 +906,7 @@ jobs:
precision: ${{ matrix.config.precision }}
router: ${{ matrix.config.router && toJson(matrix.config.router) || '' }}
kv-p2p-transfer: ${{ matrix.config['kv-p2p-transfer'] || '' }}
conc-list: ${{ toJson(matrix.config.conc) }}
conc-list: ${{ format('[{0}]', matrix.config['eval-conc']) }}
spec-decoding: ${{ matrix.config.spec-decoding }}
disagg: ${{ matrix.config.disagg }}
prefill-hardware: ${{ matrix.config.prefill.hardware }}
Expand Down
Loading
Loading