APIX / Methodology

Scores you can audit.

APIX is designed to separate measured performance from popularity. This MVP shows the scoring structure; its current tool values are demo data, not published benchmark results.

01 · Define

Task suites

Each category gets repeatable tasks with clear success criteria, inputs and expected outputs.

02 · Test

Lab evidence

Tools are tested against the same task suite, with model/version, timing, cost and outputs recorded.

03 · Validate

User evidence

Verified usage evidence can complement lab tests without replacing them.

Proposed score structure

  • Lab score: repeatable benchmark performance by category.
  • User score: structured feedback from verified product use.
  • APIX score: weighted synthesis of relevant evidence.
  • Confidence: coverage, sample size, recency and version stability.

Version-aware

AI products change quickly. Scores should belong to a tested product/model version and date, not live forever under a brand name.

Category-specific

A tool can be excellent for image generation and irrelevant for research. APIX avoids pretending one universal benchmark fits every task.

Evidence-linked

The production goal is for every official score to link back to prompts, outputs, evaluators, timings and costs.

MVP status: all six current product profiles and scores are illustrative demo content. They demonstrate the experience and scoring model, not claims about real products.
Evidence integrity

Every important claim needs provenance.

VIQQO stores source type, observation date, verification date, freshness window, verification status and evidence level separately from the claim itself. Vendor documentation can support factual product attributes, but vendor marketing is not independent performance proof. Prototype records remain marked demo until real sources replace them.