Automatically evaluate AI models,
and see why they behave as they do
Supermagin automatically evaluates AI models, analyzes their outputs, detects failures and patterns, and gives teams actionable insights through a software platform. Connect your own models and keys — the software does the rest.
BYOK-friendly · Works with OpenAI, Anthropic, Gemini & Ollama
See the whole story
Logs, traces, outputs, failures, retries, and confidence signals. Every detail captured automatically.
Compare any models
Run the same tasks across OpenAI, Claude, Gemini, DeepSeek, Ollama, and custom models.
Live dashboards
Watch long workflows unfold with streaming metrics, live logs, and automated root-cause hints.
Automated root-cause analysis
The evaluation engine finds why a model fails, using correlation and structure.
Chief Investigator — Automated AI Evaluation Engine
Chief Investigator is Supermagin's automated AI evaluation engine. It automatically analyzes AI model outputs, identifies failures and behavioral patterns, generates evaluation diagnostics, and surfaces findings to your team inside Supermagin. No manual review required — the engine runs on the software platform.
Run
The platform runs your configured models on your evaluation tasks.
Analyze
Chief Investigator automatically processes the outputs and metrics.
Findings
Diagnostics, patterns, and root causes appear in your workspace.
How it works
Three steps to start evaluating your models.
Connect your stack
Link your models, datasets, databases, and tools. BYOK means you control API spend.
Define tasks
Create evaluation tasks for apps, games, simulators, websites, or any workflow.
Run and observe
Run multiple models on the same tasks. Watch logs, outputs, and metrics stream live, then read the Chief Investigator’s automated verdict.
For AI engineers
Debug faster. Ship safer.
- See exactly where prompts, tools, or data cause failures.
- Drill down from a failing task to the specific step and log line.
- Catch hallucinations, brittle workflows, and regressions before production.
For companies and labs
Compare vendors. Build a trusted quality layer.
- Run the same benchmarks across multiple providers and internal models.
- Know which model actually wins for your specific tasks and data.
- Standardize evaluation across teams with shared dashboards and audit logs.
Audit any website, then get an AI optimization plan
Enter a URL and the software crawls the site and runs real SEO, GEO and AEO audits, reliability and security checks, and math diagnostics. Connect your own AI model and it produces an automated visibility and ranking assessment with a step-by-step optimization plan.
- Real crawl: every resource, header, and metric measured by the software
- SEO · GEO · AEO scores with keyword coverage and an optimization kit
- Bring your own AI model for an automated visibility and ranking assessment
- White-label client reports in the Website Optimizer
Software-generated analysis. No manual services. Not a ranking guarantee.
Crawl & audit
Every CSS, JS, image, font, header, and security check the page actually loads, with real status, size, and load time.
SEO · GEO · AEO scores
Deterministic audits score how findable the site is for your target keywords and how well it answers AI engines.
AI optimization plan
Connect your own model and the software synthesizes a visibility score, prioritized fixes, and a ranking plan from the real crawl data.
Control your costs
You own your keys. The software handles orchestration and observability.
Bring your own key
Use your own OpenAI, Claude, Gemini, DeepSeek, and Ollama keys. You pay providers directly for inference. The software orchestrates and observes.
Simple subscriptions
Recurring software subscriptions billed via Polar for global payments. Starter $19 · Pro $49 · Team $199 per month. Cancel anytime.
Plans for every team
Software subscriptions for automated AI evaluation. BYOK, so you keep your own keys.
Starter
Individual AI evaluation. Automated runs, basic diagnostics, workspace dashboard.
Pro
Advanced automated diagnostics, model comparison, and API access.
Team
Team workspaces, collaboration, RBAC, webhooks, and organizational controls.
Starter, Pro, and Team are recurring software subscriptions for access to the Supermagin platform. Polar will be used solely to process payments for these three software subscriptions. Customers operate the platform themselves; evaluation and analysis are performed by the software. No consulting or manual professional services are included.
Start evaluating your models today
Connect your own keys and tasks, and let the software run your evaluations.