Skip to content
llmwaves

Compare

Four models, one page

Pick the models you are actually choosing between. Every row is scaled within itself, so a bar means something next to its neighbours and nothing across rows.

SpecificationWeaver (alpha)MancerClaude Opus 5Anthropic
ReleasedAug 2, 2023Jul 24, 2026
Context window8K1M
Max output2K128K
Input / 1M$0.5$5
Output / 1M$0.75$25
Cost per task$0.04$2.25
Arena Elo1,393
Serving providers15
ParametersUndisclosedUndisclosed
LicenceProprietaryProprietary
Capabilities
    • Reasoning
    • Tool use
    • Vision

    Metrics side by side

    Each row is scaled to the largest value in that row — bars compare within a row, never across rows

    • Weaver (alpha)
    • Claude Opus 5
    • Intelligence Index

      Weaver (alpha)not measured
      Claude Opus 5
      63.1
    • Coding Index

      Weaver (alpha)not measured
      Claude Opus 5
      78.0
    • Agentic Index

      Weaver (alpha)not measured
      Claude Opus 5
      59.2
    • Output speed

      Weaver (alpha)not measured
      Claude Opus 5
      87 t/s
    • Context window

      Weaver (alpha)
      8K
      Claude Opus 5
      1M
    • Latency · lower is better

      Weaver (alpha)not measured
      Claude Opus 5
      1.12s
    • Blended price / 1M · lower is better

      Weaver (alpha)
      $0.563
      Claude Opus 5
      $10
    • Cost per task · lower is better

      Weaver (alpha)
      $0.04
      Claude Opus 5
      $2.25

    Rows marked “lower is better” still draw a longer bar for a larger number — read the value, not just the length. Arena Elo is in the specification table above instead: it has no meaningful zero, so a bar would flatten the gaps.

    View as table
    MetricWeaver (alpha)Claude Opus 5
    Intelligence Index63.1
    Coding Index78.0
    Agentic Index59.2
    Output speed87 t/s
    Context window8K1M
    Latency · lower is better1.12s
    Blended price / 1M · lower is better$0.563$10
    Cost per task · lower is better$0.04$2.25

    Evaluation scores

    Percentage correct on a common 0–100% scale

    • Weaver (alpha)
    • Claude Opus 5
    • GPQA Diamond

      Weaver (alpha)not measured
      Claude Opus 5
      93.2%
    • Humanity's Last Exam

      Weaver (alpha)not measured
      Claude Opus 5
      54.9%
    • SciCode

      Weaver (alpha)not measured
      Claude Opus 5
      55.7%
    • τ²-bench

      Weaver (alpha)not measured
      Claude Opus 5
      42.1%
    • AA-LCR long context

      Weaver (alpha)not measured
      Claude Opus 5
      75.7%

    A missing bar means that evaluation was not run for that model — it is not a zero.

    View as table
    EvaluationWeaver (alpha)Claude Opus 5
    GPQA Diamond93.2%
    Humanity's Last Exam54.9%
    SciCode55.7%
    τ²-bench42.1%
    Terminal-Bench Hard
    LiveCodeBench
    AA-LCR long context75.7%
    AIME 2025

    Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).