# Cohere vs Together AI

**Leader by Vioscale score:** Cohere

| Attribute | Cohere | Together AI |
|---|---|---|
| **Vioscale score** | 79.3 (45% (low)) | 72.8 (60% (medium)) |
| deployment.options | `{"cloud":true,"hybrid":true,"on_prem":true,"self_hosted":true}` | `{"cloud":true}` |
| description.long | A suite of AI tools including high-performance language models, semantic search and discovery, speech-to-text, and embedding services. Designed to run securely on your own infrastructure or via a managed cloud service, with options to train custom models on proprietary data. | Together AI provides cloud infrastructure for hosting and serving open-source language, image, audio, and video models through serverless inference, reserved capacity, and dedicated GPU instances. The platform includes fine-tuning capabilities, batch processing, managed storage, and GPU cluster support for custom model development and training—eliminating the need for users to manage underlying infrastructure. |
| features.capabilities | `{"open_source":false,"soc2_type_ii":true,"compute_model":"multi_tenant_hybrid","pricing_model":"usage-based-compute","batch_processing":false,"deployment_model":"both","data_residency_eu":true,"real_time_latency":true,"embeddings_endpoint":true,"custom_model_training":true,"domain_specific_models":true,"openai_compatible_rest_api_schema":false,"bring_your_own_weights_byow_hosting":false,"dynamic_batching_and_kv_cache_management":false,"massive_context_window_support_100k_plus":false,"hipaa_baa_compliance_for_medical_inference":false,"multimodal_vision_and_audio_in_out_support":false,"native_function_calling_and_strict_json_mode":false,"streaming_server_sent_events_sse_token_yield":false,"lpu_or_custom_silicon_for_extreme_low_latency":false,"zero_data_retention_training_opt_out_enterprise":true}` | `{"role":"platform","evaluation":false,"soc2_type_ii":false,"compute_model":"multi_tenant_hybrid","managed_cloud":true,"model_serving":true,"pricing_model":"per_million_tokens","self_hostable":false,"multi_provider":true,"vpc_deployment":true,"otel_compatible":false,"no_train_on_data":"yes","llm_observability":false,"prompt_management":false,"framework_agnostic":true,"experiment_tracking":false,"openai_compatible_rest_api_schema":false,"bring_your_own_weights_byow_hosting":true,"dynamic_batching_and_kv_cache_management":false,"massive_context_window_support_100k_plus":false,"hipaa_baa_compliance_for_medical_inference":false,"multimodal_vision_and_audio_in_out_support":true,"native_function_calling_and_strict_json_mode":false,"streaming_server_sent_events_sse_token_yield":false,"lpu_or_custom_silicon_for_extreme_low_latency":false,"zero_data_retention_training_opt_out_enterprise":true}` |
| market.availability | `{"hqCountry":"CA","primaryMarkets":["US","CA"],"availabilityScope":"global","availableCountries":[],"notAvailableCountries":[]}` | `{"primaryMarkets":["US"],"availabilityScope":"global","availableCountries":[],"notAvailableCountries":[]}` |
| platform.support | `{"cli":true,"web":true}` | `{"cli":true,"web":true}` |
| pricing | `{"type":"hybrid","plans":[{"free":false,"name":"API - Command","summary":"$1.00–$2.00 per 1M tokens","commitment":"monthly","components":[{"per":{"qty":1000000,"unit":"tokens"},"kind":"metered","unit":"input","amount":1,"currency":"USD"},{"per":{"qty":1000000,"unit":"tokens"},"kind":"metered","unit":"output","amount":2,"currency":"USD"}],"contactSales":false},{"free":false,"name":"API - Command-light","summary":"$0.30–$0.60 per 1M tokens","commitment":"monthly","components":[{"per":{"qty":1000000,"unit":"tokens"},"kind":"metered","unit":"input","amount":0.3,"currency":"USD"},{"per":{"qty":1000000,"unit":"tokens"},"kind":"metered","unit":"output","amount":0.6,"currency":"USD"}],"contactSales":false},{"free":false,"name":"API - Command R","summary":"$0.50–$1.50 per 1M tokens","commitment":"monthly","components":[{"per":{"qty":1000000,"unit":"tokens"},"kind":"metered","unit":"input","amount":0.5,"currency":"USD"},{"per":{"qty":1000000,"unit":"tokens"},"kind":"metered","unit":"output","amount":1.5,"currency":"USD"}],"contactSales":false},{"free":false,"name":"API - Command R+ (latest)","summary":"$2.50–$10.00 per 1M tokens","commitment":"monthly","components":[{"per":{"qty":1000000,"unit":"tokens"},"kind":"metered","unit":"input","amount":2.5,"currency":"USD"},{"per":{"qty":1000000,"unit":"tokens"},"kind":"metered","unit":"output","amount":10,"currency":"USD"}],"contactSales":false},{"free":false,"name":"Model Vault","summary":"Pricing per instance based on model and performance tier","commitment":"monthly","description":"Dedicated, managed instance with no shared resources","contactSales":true},{"free":false,"name":"Enterprise","summary":"Custom pricing based on requirements","description":"Custom models, private deployment, and personalized support","contactSales":true}],"summary":"Free trial tier; pay-per-token from $0.30–$10.00 per 1M tokens depending on model","currency":"USD","freeTier":true,"sourceUrl":"https://cohere.com/fr/pricing","retrievedAt":"2026-08-21T13:44:41.876Z","startingPrice":{"unit":"input","amount":0.3,"currency":"USD"},"billingPeriods":["month"]}` | `{"type":"usage","plans":[{"free":false,"name":"Serverless Inference","summary":"From $0.00014–$15/1M tokens depending on model. Pay only for usage.","features":["50+ open-source models","Variable pricing by model and token type","Batch API support","Private endpoints"],"components":[{"per":{"qty":1000000,"unit":"tokens"},"kind":"metered","amount":0.00014,"currency":"USD"}],"contactSales":false},{"free":false,"name":"Provisioned Throughput","summary":"Reserved throughput capacity with 99% SLA. PTU-based pricing structure.","features":["Reserved token capacity","99% uptime SLA","Token-based pricing model","Production-grade reliability"],"commitment":"monthly","contactSales":false},{"free":false,"name":"Dedicated Inference","summary":"$3.69–$8.99 per GPU per hour (on-demand); reserved discounts available.","features":["Single-tenant GPU instances","Guaranteed performance (no resource sharing)","Custom model support","Autoscaling for traffic spikes"],"commitment":"monthly","components":[{"kind":"fixed","amount":3.69,"period":"hour","currency":"USD"}],"contactSales":false}],"addOns":[{"name":"Fine-Tuning","components":[{"per":{"qty":1000000,"unit":"tokens"},"kind":"metered","amount":0.48,"currency":"USD"}]},{"name":"Managed Storage"},{"name":"GPU Clusters"}],"summary":"Usage-based pricing starting from $0.00014/1M tokens for serverless inference. Reserved capacity and dedicated GPU instances available; on-demand GPU pricing from $3.69/hour.","currency":"USD","freeTier":false,"sourceUrl":"https://www.together.ai/pricing","retrievedAt":"2026-08-21T12:44:34.463Z","startingPrice":{"unit":"tokens","amount":0.00014,"currency":"USD"},"billingPeriods":["month"]}` |
| pricing.free_tier | yes | no |
| pricing.model | freemium | commercial |
| pricing.price_level | low | low |
| pricing.starting_price | `{"amount":0.3,"currency":"USD"}` | `{"amount":0.00014,"currency":"USD"}` |
| pricing.transparent | yes | yes |
| reliability.sla_pct | - | 99 |
| reliability.status_page | - | yes |
| security.disclosure_policy | yes | yes |
| security.gdpr | yes | yes |
| security.iso27001 | - | yes |
| security.soc2 | yes | yes |

## Capabilities (Inference Providers)

| Capability | Cohere | Together AI |
|---|:--:|:--:|
| **Capabilities** |  |  |
| Compute model | Multi tenant hybrid | Multi tenant hybrid |
| Zero data retention training opt out enterprise | ✓ | ✓ |
| Openai compatible rest API schema | ✗ | ✗ |
| Native function calling and strict json mode | ✗ | ✗ |
| Multimodal vision and audio in out support | ✗ | ✓ |
| Streaming server sent events sse token yield | ✗ | ✗ |
| Massive context window support 100k plus | ✗ | ✗ |
| Lpu or custom silicon for extreme low latency | ✗ | ✗ |
| Bring your own weights byow hosting | ✗ | ✓ |
| Dynamic batching and kv cache management | ✗ | ✗ |
| SOC2 type ii | ✓ | ✗ |
| HIPAA baa compliance for medical inference | ✗ | ✗ |
| Pricing model | usage-based-compute | Per million tokens |

*Source: Vioscale. Generated 2026-09-01T16:27:12.996Z. "-" = undocumented, not absent.*
