Module 2 Assessment | SA Enablement 2026 | Xinzhi Sherry Zhu (zhxinzhi)
| Model | Provider | Inference Profile | Characteristics |
|---|---|---|---|
| Claude Sonnet 4.6 | Anthropic | jp.anthropic.claude-sonnet-4-6 | High capability, extended thinking |
| Nova Pro | Amazon | apac.amazon.nova-pro-v1:0 | Balanced performance/cost |
| GPT-OSS 20B | OpenAI | openai.gpt-oss-20b-1:0 | Cost-optimized reasoning |
All models accessed via Amazon Bedrock Converse API with cross-region inference profiles.
Three prompt categories tested: short factual Q&A, analytical comparison, and structured summarization.
| Model | Short Factual | Analytical | Structured | Average |
|---|---|---|---|---|
| Claude Sonnet 4.6 | 3,189 | 28,519 | 5,178 | 12,295 |
| Nova Pro | 659 | 2,300 | 1,140 | 1,366 |
| GPT-OSS 20B | 558 | 6,942 | 1,593 | 3,031 |
| Model | Short Factual | Analytical | Structured |
|---|---|---|---|
| Claude Sonnet 4.6 | 87 | 1,713 | 202 |
| Nova Pro | 66 | 541 | 210 |
| GPT-OSS 20B | 93 | 2,048* | 469 |
* GPT-OSS 20B hit the default max_tokens limit (2048) on the analytical prompt.
| Model | Short Factual | Analytical | Structured | Total (3 calls) |
|---|---|---|---|---|
| Claude Sonnet 4.6 | $0.00137 | $0.02576 | $0.00310 | $0.03023 |
| Nova Pro | $0.00022 | $0.00174 | $0.00068 | $0.00264 |
| GPT-OSS 20B | $0.00005 | $0.00083 | $0.00020 | $0.00107 |
For startup chat applications: Use Nova Pro as the default model (low latency, reasonable cost), offer Claude Sonnet 4.6 as a premium option for complex tasks, and use GPT-OSS 20B for background/batch operations where cost dominates.
| Priority | Recommended Model | Rationale |
|---|---|---|
| Quality first | Claude Sonnet 4.6 | Superior reasoning, best for customer-facing |
| Latency first | Nova Pro | Sub-second for short queries, consistent |
| Cost first | GPT-OSS 20B | 28x cheaper than Claude, acceptable quality |
| Balanced | Nova Pro | Best latency/cost ratio for interactive use |