Amazon Web Services has enabled in-region inference for Anthropic’s Claude Opus 5 and Claude Sonnet 5 models in Seoul and Singapore.

The deployment routes AI workloads into the Asia Pacific (Seoul) [ap-northeast-2] and Asia Pacific (Singapore) [ap-southeast-1] regions to satisfy localized data handling mandates. Entities in financial services, healthcare, and public administration requiring single-Region data retention can now execute these models via Amazon Bedrock without traffic leaving the selected jurisdictional boundary AWS Machine Learning Blog.

Earlier Bedrock implementations used cross-Region inference profiles that depended on a centralized routing mechanism. The revised configuration removes that abstraction, compelling the bedrock-runtime endpoint to process both input prompts and generated responses strictly within one Region. This structural modification guarantees end-to-end data containment but transfers latency management and horizontal scaling obligations to the host site’s available resources.

Clients invoking Claude Opus 5 and Claude Sonnet 5 in Seoul, or Claude Sonnet 5 in Singapore, interact with the systems using the Converse API, InvokeModel API, and Anthropic Messages API. Official documentation covers setup procedures for both the Amazon Bedrock console and custom code paths. Cost accounting uses standard on-demand rates tied to the invoked Region. Telemetry dashboards, CloudWatch usage counters, and CloudTrail activity records map solely to the calling location, which simplifies monitoring by removing origin-and-destination reconciliation from cost tracking AWS Machine Learning Blog.

Throughput for these endpoints is bounded by the compute capacity of the selected Region, and all requests remain subject to per-Region service quotas AWS Machine Learning Blog. The company’s disclosure provides no external penetration testing data or third-party certification validating network isolation guarantees for these endpoints.

Engineering teams embedding these interfaces into established pipelines should tune backoff algorithms to address localized quota exhaustion rather than expecting downstream failover. Organizations running dense, continuous inference streams will require tight visibility into per-Region buffer saturation during migration. AWS has not defined deployment timelines for other territories or outlined minimum volume commitments for clients requesting quota increases past baseline on-demand allocations.