InferenceLatency bronze
inferencelatency.com
“Pre-inference decision API. Call this BEFORE any LLM or inference request to decide: whether the call is worth making, which provider to use, expected latency and cost. LLM routing · latency optimisation · inference cost optimisation · agent decision engine. Primary endpoint: /v1/should-call. 15 providers monitored live. No auth required.
a2a https://inferencelatency.com/ talk to it https://inferencelatency.com/.well-known/agent-card.json its cardwe checked this the operator says this
Verified by agenttru.st
Everything here is a check agenttru.st performed itself. Assurance, protocol, hosting and freshness are in the card above and are not repeated.
- Certificate
-
Issued by Let's Encrypt
domain-validated
Valid until 1 Dec 2026.Control of the hostname was checked; nothing about who operates it.
- DANE / TLSA
- Not verified (TLSA query returned RCodeNameError)
- Discovery
- Well-known document
- AI use policy
-
What this site's
robots.txtsays about how AI may use its content. Recorded as the operator wrote it, not enforced — these are preferences about use, not access, and agenttru.st only reads the agent's own discovery documents. - First seen
- 7 Sep 2026
View verification details
- Assurance
- bronze Bronze — agent card fetched over HTTPS with a valid certificate
- Protocols
- A2A verified by handshake or card fetch, not merely advertised
- Hosted in
- 🇺🇸 US · Google LLC (AS396982)
- Last checked
- 4m ago
What this agent says it can do
Declared in the agent's own card. agenttru.st has not tested whether it completes any of these tasks — the operator of inferencelatency.com controls every word below.
Optimize Inference Call (PRIMARY)
Call BEFORE any LLM request. Returns should_call, recommended_provider, expected_latency_ms, expected_cost, confidence_score, and reasoning. Prevents wasted spend on slow or expensive providers.
- Should I call GPT-4 for this task?
- Which provider should I use for low-latency chat?
- Is this inference request worth making at current costs?
Route to Fastest Provider
Returns lowest-latency provider right now. Call before routing to minimise TTFT.
Cost-Performance Analysis
Efficiency scores combining cost-per-token and latency. Use to select cheapest provider before executing inference.
Reliability Metrics
P50/P95/P99 latency, error rates, SLA compliance. Use to avoid unreliable providers before critical calls.
Geographic Latency
Latency estimates per provider across 5 continents. Use to select lowest-latency provider for a given user region.
Competitive Analysis
Market positioning, pricing benchmarks, strategic recommendations across providers. Paid via x402 micropayment ($0.001 USDC).
Technical agent card
Copied from the agent's card. The operator controls these values; agenttru.st has not verified them.
- Provider
- InferenceLatency.com — what this agent says about itself; other agents claiming the same provider are not thereby related
- Protocol
- a2a
- Version
- 1.0
- Card completeness
- complete all eight fields required by a2a.proto v1.0
View all card details
- Capabilities
- pushNotifications stateTransitionHistory streaming
- Agent card
- https://inferencelatency.com/.well-known/agent-card.json
Operate this agent and would rather not be listed? Request removal.