# Amazon SageMaker Inference vs LMDeploy

| Attribute | Amazon SageMaker Inference | LMDeploy |
|---|---|---|
| **Vioscale score** | 36.1 (40% (low)) | 43.8 (23% (low)) |
| activity.commits_last_30d | - | 66 |
| adoption.dependent_repos | - | 2 |
| adoption.github_stars | - | 8,023 |
| deployment.options | `{"cloud":true}` | `{"self_hosted":true}` |
| description.long | AWS SageMaker Inference is a managed service for deploying trained machine learning models into production, supporting multiple ML frameworks and providing integration with AWS MLOps tools like model registries, feature stores, and CI/CD pipelines. | A software framework that enables developers to compress, deploy, and serve large language models with quantization optimization, multiple inference engines, and compatibility across various model architectures. |
| features.capabilities | `{"gpu_acceleration":true,"multi_model_serving":true,"multi_framework_support":true,"autoscaling_scale_to_zero":true}` | - |
| integrations.count | 13 | 2 |
| integrations.list | `[{"name":"TensorFlow"},{"name":"PyTorch"},{"name":"ONNX"},{"name":"XGBoost"},{"name":"SageMaker Pipelines"},{"name":"SageMaker Projects"},{"name":"SageMaker Feature Store"},{"name":"SageMaker Model Registry"},{"name":"SageMaker Clarify"},{"name":"Amazon Bedrock"},{"name":"Amazon S3"},{"name":"AWS CloudWatch"},{"name":"AWS CloudTrail"}]` | `[{"name":"llm-compressor"},{"name":"OpenCompass"}]` |
| language.primary | - | Python |
| license.spdx | - | Apache-2.0 |
| market.availability | `{"hqCountry":"US","primaryMarkets":["US"],"availabilityScope":"global","availableCountries":["US"],"notAvailableCountries":[]}` | - |
| platform.support | `{"cli":true,"web":true}` | `{"cli":true,"windows":true}` |
| pricing | - | `{"type":"open_source","freeTier":true,"sourceUrl":"https://lmdeploy.readthedocs.io/en/latest/","retrievedAt":"2026-08-20T12:17:08.568Z"}` |
| pricing.free_tier | yes | yes |
| pricing.model | - | open_source |
| pricing.price_level | - | free |
| release.cadence_days | - | 21 |
| release.history | - | `[{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.16.0","date":"2026-08-19T04:43:53Z","type":"stable","version":"v0.16.0"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.15.0","date":"2026-07-31T13:00:46Z","type":"stable","version":"v0.15.0"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.14.0","date":"2026-06-24T04:36:56Z","type":"stable","version":"v0.14.0"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/0.14.0a2","date":"2026-06-16T04:16:04Z","type":"prerelease","version":"0.14.0a2"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/0.14.0a1","date":"2026-06-01T08:46:02Z","type":"prerelease","version":"0.14.0a1"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.13.0","date":"2026-05-12T03:46:57Z","type":"stable","version":"v0.13.0"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.12.3","date":"2026-04-08T03:37:26Z","type":"stable","version":"v0.12.3"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.12.2","date":"2026-03-18T03:13:55Z","type":"stable","version":"v0.12.2"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.12.1","date":"2026-02-13T09:02:12Z","type":"stable","version":"v0.12.1"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.12.0","date":"2026-02-04T06:28:25Z","type":"stable","version":"v0.12.0"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.11.1","date":"2025-12-24T13:27:06Z","type":"stable","version":"v0.11.1"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.11.0","date":"2025-12-04T06:20:38Z","type":"stable","version":"v0.11.0"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.10.2","date":"2025-10-28T11:32:42Z","type":"stable","version":"v0.10.2"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.10.1","date":"2025-09-26T02:45:30Z","type":"stable","version":"v0.10.1"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.10.0","date":"2025-09-09T05:05:09Z","type":"stable","version":"v0.10.0"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.9.2.post1","date":"2025-08-19T09:44:41Z","type":"stable","version":"v0.9.2.post1"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.9.2","date":"2025-07-26T10:00:25Z","type":"stable","version":"v0.9.2"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.9.1","date":"2025-07-04T10:05:31Z","type":"stable","version":"v0.9.1"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.9.0","date":"2025-06-19T02:27:57Z","type":"stable","version":"v0.9.0"},{"url":"https://github.com/InternLM/lmdeploy/releases/tag/v0.8.0","date":"2025-05-04T03:17:23Z","type":"stable","version":"v0.8.0"}]` |
| security.fedramp | yes | - |
| security.gdpr | yes | - |
| security.hipaa | yes | - |
| security.pci | yes | - |
| security.vulnerabilities | - | `{"count":6,"source":"https://advisories.ecosyste.ms/api/v1/advisories?ecosystem=pypi&package_name=lmdeploy&per_page=100","last_12m":4,"max_severity":"HIGH"}` |

## Capabilities (Model Serving)

| Capability | Amazon SageMaker Inference | LMDeploy |
|---|:--:|:--:|
| **Capabilities** |  |  |
| Continuous batching | - | - |
| Dynamic batching | - | - |
| Multi framework support | ✓ | - |
| GPU acceleration | ✓ | - |
| Multi GPU multi node | - | - |
| Quantization support | - | - |
| Openai compatible API | - | - |
| Autoscaling scale to zero | ✓ | - |
| Multi model serving | ✓ | - |
| Canary ab rollout | - | - |
| Kubernetes native | - | - |
| Open source | - | - |

*Source: Vioscale. Generated 2026-09-01T17:20:00.089Z. "-" = undocumented, not absent.*
