Skip to content
llmwaves

Compare

Four models, one page

Pick the models you are actually choosing between. Every row is scaled within itself, so a bar means something next to its neighbours and nothing across rows.

SpecificationAion-RP 1.0 (8B)AionLabsClaude Opus 5Anthropic
ReleasedFeb 4, 2025Jul 24, 2026
Context window33K1M
Max output33K128K
Input / 1M$0.8$5
Output / 1M$1.6$25
Cost per task$0.08$2.25
Arena Elo1,393
Serving providers15
ParametersUndisclosedUndisclosed
LicenceProprietaryProprietary
Capabilities
    • Reasoning
    • Tool use
    • Vision

    Metrics side by side

    Each row is scaled to the largest value in that row — bars compare within a row, never across rows

    • Aion-RP 1.0 (8B)
    • Claude Opus 5
    • Intelligence Index

      Aion-RP 1.0 (8B)not measured
      Claude Opus 5
      63.1
    • Coding Index

      Aion-RP 1.0 (8B)not measured
      Claude Opus 5
      78.0
    • Agentic Index

      Aion-RP 1.0 (8B)not measured
      Claude Opus 5
      59.2
    • Output speed

      Aion-RP 1.0 (8B)
      10 t/s
      Claude Opus 5
      87 t/s
    • Context window

      Aion-RP 1.0 (8B)
      33K
      Claude Opus 5
      1M
    • Latency · lower is better

      Aion-RP 1.0 (8B)
      821ms
      Claude Opus 5
      1.12s
    • Blended price / 1M · lower is better

      Aion-RP 1.0 (8B)
      $1
      Claude Opus 5
      $10
    • Cost per task · lower is better

      Aion-RP 1.0 (8B)
      $0.08
      Claude Opus 5
      $2.25

    Rows marked “lower is better” still draw a longer bar for a larger number — read the value, not just the length. Arena Elo is in the specification table above instead: it has no meaningful zero, so a bar would flatten the gaps.

    View as table
    MetricAion-RP 1.0 (8B)Claude Opus 5
    Intelligence Index63.1
    Coding Index78.0
    Agentic Index59.2
    Output speed10 t/s87 t/s
    Context window33K1M
    Latency · lower is better821ms1.12s
    Blended price / 1M · lower is better$1$10
    Cost per task · lower is better$0.08$2.25

    Evaluation scores

    Percentage correct on a common 0–100% scale

    • Aion-RP 1.0 (8B)
    • Claude Opus 5
    • GPQA Diamond

      Aion-RP 1.0 (8B)not measured
      Claude Opus 5
      93.2%
    • Humanity's Last Exam

      Aion-RP 1.0 (8B)not measured
      Claude Opus 5
      54.9%
    • SciCode

      Aion-RP 1.0 (8B)not measured
      Claude Opus 5
      55.7%
    • τ²-bench

      Aion-RP 1.0 (8B)not measured
      Claude Opus 5
      42.1%
    • AA-LCR long context

      Aion-RP 1.0 (8B)not measured
      Claude Opus 5
      75.7%

    A missing bar means that evaluation was not run for that model — it is not a zero.

    View as table
    EvaluationAion-RP 1.0 (8B)Claude Opus 5
    GPQA Diamond93.2%
    Humanity's Last Exam54.9%
    SciCode55.7%
    τ²-bench42.1%
    Terminal-Bench Hard
    LiveCodeBench
    AA-LCR long context75.7%
    AIME 2025

    Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).