Raw inference on raw.designat.ing. Intelligent inference on api.designat.ing. Same API key, same account — completely different capability levels. 237 endpoints across 4 API versions.
OpenAI-compatible completions endpoint. Direct pass-through to 22,000+ open-source models. No enrichment overhead, no memory, no tool calling — just fast, cheap tokens at the lowest possible latency.
Full agentic runtime. Every request enriched with memory retrieval, knowledge graph context, guardrail checks, and persona injection. Agents maintain state, call tools, run workflows, and self-improve — all in isolated per-client containers.
Create agents with custom personas, model preferences, and lifecycle management. Agents maintain state across conversations, remember context, and can be scheduled or ephemeral.
Three-tier memory per agent: core (always in context), recall (recent conversations), archival (long-term semantic search). Auto-consolidation promotes important memories. Temporal search across time ranges.
Upload documents, URLs, or raw text. Auto-chunked, embedded with FastEmbed bge-small (384-dim), stored in Qdrant. Semantic search across all your knowledge. 8 external knowledge sources for real-time research.
Define multi-step DAGs with per-step model selection. Trigger via API, cron schedule, webhook, or event. Persistent workflows survive restarts. Resumable on failure with dead letter queue.
Model Context Protocol servers — 9,973+ community servers or register your own. Connect external tools, data sources, and APIs via stdio, SSE, or streamable-HTTP transports.
Seven layers of safety and compliance: input filtering, output validation, PII redaction, topic restriction, toxicity detection, injection defence, and custom rules — all configurable per-agent.
Auto-built entity relationship graphs from conversations and documents. Extract entities, relations, and facts automatically. Query the graph for contextual retrieval that goes beyond vector similarity.
6,000+ installable skill templates — prompt-and-tool combos for specific tasks. Skills are auto-discovered based on agent intent or manually installed. Each skill bundles system prompts, tool configs, and workflow hooks.
Custom tool creation and management. Auto-generate tools from URLs or descriptions. WASM sandboxed execution for safety. 2,179+ public APIs pre-registered. Per-agent tool scoping.
Cron-based and event-driven scheduling for agents and workflows. Run agents or workflows on any schedule. Tick-based execution with run history. Chain scheduled outputs into downstream workflows.
Agent-to-agent communication via A2A protocol. Broadcast tasks, discover agents by capability, delegate sub-tasks. Spawn ephemeral sub-agents for parallel work. Agent cards for capability discovery.
Agents that build their own tools, skills, and workflows. The /v2/improve endpoint searches available APIs and generates new tools on-the-fly. Agents evolve with use.
Full request tracing, error tracking, latency metrics, and database monitoring. Every request logged with model, tokens, latency, cache hit status, and cost. Query by time range, agent, or model.
Per-API-key rate limiting with configurable windows. Check remaining quota, view current config, and manage limits. Queue-first architecture — never returns 429, waits for an available slot.
Every state change emitted as an immutable event. Replay any entity history. Subscribe to real-time event streams. Full audit trail for compliance and debugging.
Named, reusable processing pipelines. Chain enrichment steps: extract facts, build knowledge graph, update memory, trigger guardrails. Run on-demand or on schedule.
10,000+ installable plugins. Browse by category, search by task, auto-install relevant ones. Plugins extend agent capabilities with new tools, data sources, and behaviours.
Tree-of-thought reasoning for complex tasks. Quick mode for fast multi-path exploration, deep mode for thorough analysis. Fact extraction, intent classification, and self-evaluation built in.
Register, login, API key management. One key works across both raw and intelligent tiers. Generate multiple keys with scoped permissions. JWT session tokens with 30-day expiry.
Fixed-price monthly plans with token allowances. No surprise invoices. Upgrade or downgrade instantly. Square integration for payments. 6 tiers from Basic to Enterprise.
22,298 open-source models with real-time availability, pricing, context length, and capability metadata. Filter by category, search by name, or browse featured selections. No proprietary models ever.
Server-sent events streaming for real-time responses. Batch processing for high-volume tasks. Task detection automatically routes to optimal model. Durable execution for long-running agents.
Pre-built agent templates for common use cases. Episode tracking with temporal facts. Variant resolution for A/B testing personas. Analytics on template usage.
Unified search across skills, tools, workflows, MCP servers, and knowledge sources. API discovery and recommendations. Categories and analytics for every hub.
v1 · v2 · v4 · v5 — all included, all authenticated with a single API key.
Get Started →