PromptRelay is currently in Public Beta (v0.4). Core routing and semantic caching are 100% free for early developer cohorts.
Transparent Pricing for Universal Multi-Provider AI Gateway (BYOK)
Zero setup fees during Public Beta v0.4 for PromptRelay: Universal Multi-Provider AI Gateway & AWS Bedrock SigV4 Bridge (BYOK). Bring Your Own Key (BYOK) with unlimited model routing and single-tenant VPC deployments.
Developer Beta
ACTIVE NOWFull-featured access to the Universal Multi-Provider AI Gateway, AWS SigV4 bridge, and Aurora pgvector semantic cache.
- 10,000,000 tokens / month
- Sub-20ms semantic caching
- AWS SigV4 in-memory translation
- Zero-downtime Bedrock failover
- Standard community & Discord support
- Zero-Prompt Persistence guarantee
Production Team
POST-BETA PREVIEWDedicated rate limits, team token FinOps, and multi-region priority queues.
- 100,000,000 tokens / month
- Multi-tenant team cost attribution
- Custom cache similarity thresholds
- Priority low-latency relay queues
- Automated runaway prompt throttling
- Priority email & Slack support
Dedicated VPC
ENTERPRISESelf-hosted ECS Fargate proxy deployed directly inside your private AWS VPC.
- Unlimited token throughput
- Self-hosted AWS ECS proxy relay
- AWS VPC PrivateLink endpoints
- Custom HIPAA BAA & SOC2 addendum
- 99.99% availability SLA with financial credits
- Dedicated Solutions Architect support
Detailed Tier Comparison
Every tier guarantees Zero-Prompt Persistence and AWS SigV4 protocol support.
| Feature | Developer Beta ($0) | Production Team ($49) | Dedicated VPC (Custom) |
|---|---|---|---|
| Monthly Token Quota | 10M tokens | 100M tokens | Custom / Unlimited |
| AWS SigV4 Bedrock Bridge | Included | Included | Included |
| Aurora pgvector Semantic Cache | Shared Partition | Dedicated Partition | Customer VPC Aurora |
| Stream-Teeing Latency Overhead | < 1ms | < 0.5ms | < 0.2ms |
| Zero-Prompt Persistence | Volatile RAM only | Volatile RAM only | Volatile RAM only |
| AWS KMS Envelope Encryption | KMS Managed | Tenant-Scoped KMS | Customer-Owned CMK |
| Uptime SLA Commitment | Best Effort | 99.9% SLA | 99.99% Financial SLA |
| Support Channels | Discord & Community | Email & Shared Slack | 24/7 Dedicated AWS SA |
Frequently Asked Questions
Estimate Cost Reductions for Your Monthly Inference Volume
Without semantic cache offloading
58.0M tokens served free from cache
$5,220 annualized cost reduction
Eliminated client wait-time for end users every single month.
Amazon Titan Embeddings v2 with HNSW cosine distance > 0.85.
Need Custom Throughput or On-Premise Gateway?
Direct Technical Lead Inquiry
Dedicated Routing Channels
Enterprise Sandbox
enterprise@kiyaslabs.techProvisioning for enterprise subscriptions, dedicated VPC relays, and HIPAA BAAs.
Developer Support
support@kiyaslabs.techInquiries regarding the AWS SigV4 translation bridge, OpenAI SDK drop-in, or stream-teeing.
Security Disclosures
security@kiyaslabs.techEncrypted PGP channel for vulnerability reporting, pen-test reports, and compliance reviews.
Operating Entity for PromptRelay™ • Response SLA: Within 1 Business Day