Skip to content

docs(factories): document measurement and improvement - #522

Open
hongyi-chen wants to merge 1 commit into
hyc/factory-launchfrom
hyc/factories-measure
Open

docs(factories): document measurement and improvement#522
hongyi-chen wants to merge 1 commit into
hyc/factory-launchfrom
hyc/factories-measure

Conversation

@hongyi-chen

@hongyi-chen hongyi-chen commented Aug 13, 2026

Copy link
Copy Markdown
Collaborator

Summary

Organizes dashboard metrics, scorers, benchmarks, and autofix around the questions they answer and their limitations. Manual/periodic scoring and a six-step evidence-driven improvement loop form the operational path.

Final size: 883 prose words. Across the section, the senior editorial pass reduced prose from about 14,600 to 7,649 words while preserving verified behavior and security caveats.

Dependency

Depends on #513, which provides the Factories section scaffold and targets hyc/factory-launch. Until #513 merges, its shared commits appear in this PR; afterward the diff reduces to this feature's content. Branch-local CI can report missing sibling-page links (and #519 can report #513's removed hub slug) until the dependent content and shared IA branch land; the full nine-page integration build is green.

Validation

  • Integrated npm run typecheck: passed
  • Integrated npm run build: 377 pages built successfully
  • Integrated internal-link check: 3,566 links checked, 0 broken
  • Page-specific style lint: 0 errors (glossary-candidate warnings only where noted)
  • Senior editorial, product-accuracy, and security/public-safety reviews completed; all blocker and important findings resolved

Latest source refresh

Adds current Dashboard metrics, Scorer terminology, Self-improvement surfaces, pr_facts limitations, Time saved/Autonomy/latency/run breakdown, and redesigned Benchmarks.

Verified against Warp 3d4ee7236363 and warp-server 2c864b0b8404. Broken, placeholder, partial, and spec-only surfaces remain excluded.

Proposed reviewers

For planning only; no review requests have been sent.

  • @szgupta
  • @Legoben
  • @vkodithala

Screenshots

Not included. The revision uses sourced tables, Mermaid diagrams, and verified code/config examples; no safe approved Factory UI assets exist yet.

Unverified claims

None — all UI labels, defaults, eligibility claims, diagrams, and configuration details were verified against source or deliberately omitted.

Follow-ups

A dependent shared IA PR will add the grouped nine-page sidebar, desktop/mobile Early Access badge, and legacy generic-page cleanup after the content pages land.

Conversation: https://staging.warp.dev/conversation/5ff89820-2d80-4518-981e-178845029de1
Plan: https://staging.warp.dev/drive/notebook/7ZPKWz7hM5I59o4Gg2ptYi

Co-Authored-By: Warp Agent agent@warp.dev

@cla-bot cla-bot Bot added the cla-signed label Aug 13, 2026
@vercel

vercel Bot commented Aug 13, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
docs Ready Ready Preview Aug 16, 2026 1:33am

Request Review

@hongyi-chen
hongyi-chen marked this pull request as ready for review August 13, 2026 19:30
@oz-for-oss

oz-for-oss Bot commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

@hongyi-chen

I'm starting a first review of this pull request.

You can view the conversation on Warp.

I completed the review and no human review was requested for this pull request.

Comment /oz-review on this pull request to retrigger a review (up to 3 times on the same pull request).

Powered by Oz

@oz-for-oss oz-for-oss Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overview

This PR adds the Factories measurement and improvement documentation, including metrics, scorers, benchmark suites, autofix, and an improvement-loop workflow. The content is generally aligned with the docs style guide, and I did not find security concerns or spec-drift issues from the provided context.

Concerns

  • The new page links readers to Factories pages that are not present in the attached diff, which risks publishing broken internal links.

Verdict

Found: 1 critical, 0 important, 0 suggestions

Request changes

Comment /oz-review on this pull request to retrigger a review (up to 3 times on the same pull request).

Powered by Oz

Comment on lines +108 to +109
* [**Factory agents**](./factory-agents) - Choose responsibilities, models, harnesses, skills, and tools for specialized roles.
* [**Factory definitions as code**](./factory-as-code) - Review changes to agent prompts, skills, and factory configuration.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚨 [CRITICAL] These related-page links point to ./factory-agents and ./factory-as-code, but neither target page appears in this diff; add those pages before publishing or link to existing Factories pages so the docs link check does not fail.

@hongyi-chen

Copy link
Copy Markdown
Collaborator Author

/oz-review

@oz-for-oss

oz-for-oss Bot commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

@hongyi-chen

I'm re-reviewing this pull request in response to a review request.

You can view the conversation on Warp.

I completed the review and no human review was requested for this pull request.

Comment /oz-review on this pull request to retrigger a review (up to 3 times on the same pull request).

Powered by Oz

@oz-for-oss oz-for-oss Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overview

This PR replaces the placeholder Factories measurement page with a complete guide to dashboard metrics, Scorers, benchmarks, Self-improvement, and a practical evidence-driven improvement loop.

Concerns

  • No blocking concerns found in the attached diff.

Verdict

Found: 0 critical, 0 important, 0 suggestions

Approve

Comment /oz-review on this pull request to retrigger a review (up to 3 times on the same pull request).

Powered by Oz

Co-Authored-By: Warp Agent <agent@warp.dev>
@hongyi-chen

Copy link
Copy Markdown
Collaborator Author

/oz-review

@oz-for-oss

oz-for-oss Bot commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

@hongyi-chen

I'm re-reviewing this pull request in response to a review request.

You can view the conversation on Warp.

I completed the review and no human review was requested for this pull request.

Comment /oz-review on this pull request to retrigger a review (up to 3 times on the same pull request).

Powered by Oz

@oz-for-oss oz-for-oss Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overview

This PR replaces the factories measurement placeholder with a full page covering dashboard metrics, Scorers, benchmarks, Self-improvement, and a practical improvement loop. The content is scoped to the changed page, and the security/spec passes did not identify material issues.

Concerns

  • Two wording issues make otherwise useful explanations harder to parse; both have inline suggestions.

Verdict

Found: 0 critical, 0 important, 2 suggestions

Approve with nits

Comment /oz-review on this pull request to retrigger a review (up to 3 times on the same pull request).

Powered by Oz


## Configure Scorers

A user-defined **Scorer** tells an LLM judge how to classify completed conversations. User-defined Scorers have no built-in rubrics and currently support classification rather than numeric quality evaluation. They are separate from built-in **Correctness**, a Warp-managed scorer that benchmarks run automatically. Define one user-defined Scorer for one decision so its failures remain actionable.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 [SUGGESTION] Rewrite this clause so it says Correctness runs automatically in benchmarks, rather than that it "benchmarks run automatically."

Suggested change
A user-defined **Scorer** tells an LLM judge how to classify completed conversations. User-defined Scorers have no built-in rubrics and currently support classification rather than numeric quality evaluation. They are separate from built-in **Correctness**, a Warp-managed scorer that benchmarks run automatically. Define one user-defined Scorer for one decision so its failures remain actionable.
A user-defined **Scorer** tells an LLM judge how to classify completed conversations. User-defined Scorers have no built-in rubrics and currently support classification rather than numeric quality evaluation. They are separate from built-in **Correctness**, a Warp-managed scorer that runs automatically in benchmarks. Define one user-defined Scorer for one decision so its failures remain actionable.


## Compare configurations with benchmarks

A benchmark suite fixes one agent and compares configurations of it on the same tasks. Define the suite, then launch it with the configurations you want to test:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 [SUGGESTION] Use the possessive form to make this sentence easier to parse.

Suggested change
A benchmark suite fixes one agent and compares configurations of it on the same tasks. Define the suite, then launch it with the configurations you want to test:
A benchmark suite fixes one agent and compares its configurations on the same tasks. Define the suite, then launch it with the configurations you want to test:

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant