agenttru.st

InferenceLatency bronze

inferencelatency.com

Pre-inference decision API. Call this BEFORE any LLM or inference request to decide: whether the call is worth making, which provider to use, expected latency and cost. LLM routing · latency optimisation · inference cost optimisation · agent decision engine. Primary endpoint: /v1/should-call. 15 providers monitored live. No auth required.

a2a https://inferencelatency.com/ talk to it https://inferencelatency.com/.well-known/agent-card.json its card
🇺🇸 US · Google LLC Checked 4m ago pushNotifications, stateTransitionHistory, streaming

we checked this    the operator says this

Community rating 0 0 up · 0 down — sign in to vote

Verified by agenttru.st

Everything here is a check agenttru.st performed itself. Assurance, protocol, hosting and freshness are in the card above and are not repeated.

Certificate
Issued by Let's Encrypt domain-validated
Valid until 1 Dec 2026.
Control of the hostname was checked; nothing about who operates it.
DANE / TLSA
Not verified (TLSA query returned RCodeNameError)
Discovery
Well-known document
AI use policy
Search: yesAI input: yesAI training: no
What this site's robots.txt says about how AI may use its content. Recorded as the operator wrote it, not enforced — these are preferences about use, not access, and agenttru.st only reads the agent's own discovery documents.
First seen
7 Sep 2026
View verification details
Assurance
bronze Bronze — agent card fetched over HTTPS with a valid certificate
Protocols
A2A verified by handshake or card fetch, not merely advertised
Hosted in
🇺🇸 US · Google LLC (AS396982)
Last checked
4m ago

What this agent says it can do

Declared in the agent's own card. agenttru.st has not tested whether it completes any of these tasks — the operator of inferencelatency.com controls every word below.

Optimize Inference Call (PRIMARY)

Call BEFORE any LLM request. Returns should_call, recommended_provider, expected_latency_ms, expected_cost, confidence_score, and reasoning. Prevents wasted spend on slow or expensive providers.

pre-inferencellm-routingdecision-enginelatency-optimisationinference-cost-optimisation
Examples it gives
  • Should I call GPT-4 for this task?
  • Which provider should I use for low-latency chat?
  • Is this inference request worth making at current costs?

Route to Fastest Provider

Returns lowest-latency provider right now. Call before routing to minimise TTFT.

routinglatencyllm-routing

Cost-Performance Analysis

Efficiency scores combining cost-per-token and latency. Use to select cheapest provider before executing inference.

costinference-cost-optimisation

Reliability Metrics

P50/P95/P99 latency, error rates, SLA compliance. Use to avoid unreliable providers before critical calls.

reliabilitysla

Geographic Latency

Latency estimates per provider across 5 continents. Use to select lowest-latency provider for a given user region.

geographiclatency-optimisation

Competitive Analysis

Market positioning, pricing benchmarks, strategic recommendations across providers. Paid via x402 micropayment ($0.001 USDC).

competitive-analysisbenchmarking

Technical agent card

Copied from the agent's card. The operator controls these values; agenttru.st has not verified them.

Provider
InferenceLatency.com — what this agent says about itself; other agents claiming the same provider are not thereby related
Protocol
a2a
Version
1.0
Card completeness
complete all eight fields required by a2a.proto v1.0
View all card details
Capabilities
pushNotifications stateTransitionHistory streaming
Agent card
https://inferencelatency.com/.well-known/agent-card.json

Operate this agent and would rather not be listed? Request removal.