Skip to content
llmwaves

Compare

Four models, one page

Pick the models you are actually choosing between. Every row is scaled within itself, so a bar means something next to its neighbours and nothing across rows.

SpecificationWeaver (alpha)MancerClaude Opus 5Anthropic
ReleasedAug 2, 2023Jul 24, 2026
Context window8K1M
Max output2K128K
Input / 1M$0.5$5
Output / 1M$0.75$25
Cost per task$0.04$2.25
Arena Elo—1,393
Serving providers15
ParametersUndisclosedUndisclosed
LicenceProprietaryProprietary
Capabilities
    • Reasoning
    • Tool use
    • Vision

    Metrics side by side

    Each row is scaled to the largest value in that row — bars compare within a row, never across rows

    • Weaver (alpha)
    • Claude Opus 5
    • Intelligence Index

      Weaver (alpha)not measured
      Claude Opus 5
      63.1
    • Coding Index

      Weaver (alpha)not measured
      Claude Opus 5
      78.0
    • Agentic Index

      Weaver (alpha)not measured
      Claude Opus 5
      59.2
    • Output speed

      Weaver (alpha)not measured
      Claude Opus 5
      87 t/s
    • Context window

      Weaver (alpha)
      8K
      Claude Opus 5
      1M
    • Latency · lower is better

      Weaver (alpha)not measured
      Claude Opus 5
      1.12s
    • Blended price / 1M · lower is better

      Weaver (alpha)
      $0.563
      Claude Opus 5
      $10
    • Cost per task · lower is better

      Weaver (alpha)
      $0.04
      Claude Opus 5
      $2.25

    Rows marked “lower is better” still draw a longer bar for a larger number — read the value, not just the length. Arena Elo is in the specification table above instead: it has no meaningful zero, so a bar would flatten the gaps.

    View as table
    MetricWeaver (alpha)Claude Opus 5
    Intelligence Index—63.1
    Coding Index—78.0
    Agentic Index—59.2
    Output speed—87 t/s
    Context window8K1M
    Latency · lower is better—1.12s
    Blended price / 1M · lower is better$0.563$10
    Cost per task · lower is better$0.04$2.25

    Evaluation scores

    Percentage correct on a common 0–100% scale

    • Weaver (alpha)
    • Claude Opus 5
    • GPQA Diamond

      Weaver (alpha)not measured
      Claude Opus 5
      93.2%
    • Humanity's Last Exam

      Weaver (alpha)not measured
      Claude Opus 5
      54.9%
    • SciCode

      Weaver (alpha)not measured
      Claude Opus 5
      55.7%
    • τ²-bench

      Weaver (alpha)not measured
      Claude Opus 5
      42.1%
    • AA-LCR long context

      Weaver (alpha)not measured
      Claude Opus 5
      75.7%

    A missing bar means that evaluation was not run for that model — it is not a zero.

    View as table
    EvaluationWeaver (alpha)Claude Opus 5
    GPQA Diamond—93.2%
    Humanity's Last Exam—54.9%
    SciCode—55.7%
    τ²-bench—42.1%
    Terminal-Bench Hard——
    LiveCodeBench——
    AA-LCR long context—75.7%
    AIME 2025——

    Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).