# Apache Spark vs Census

**Leader by Vioscale score:** Census

| Attribute | Apache Spark | Census |
|---|---|---|
| **Vioscale score** | 60.5 (38% (low)) | 82.3 (74% (medium)) |
| activity.commits_last_30d | 100 | - |
| adoption.dependent_repos | 8,896 | - |
| adoption.github_stars | 43,882 | - |
| deployment.options | `{"on_prem":true,"self_hosted":true}` | `{"cloud":true,"hybrid":true}` |
| description.long | Apache Spark is an open-source, multi-language distributed computing engine that unifies data engineering, data science, and machine learning workloads. It processes data at scale using batch or streaming paradigms, provides SQL query capabilities for analytics, and includes built-in libraries for machine learning and graph processing. | A fully managed data integration platform that moves data from 900+ sources into warehouses, lakes, and applications, then transforms and activates it for analytics and AI. Supports real-time and scheduled synchronization with enterprise security and compliance certifications. |
| features.capabilities | `{"role":"engine","paradigm":"both","self_hosted_oss":true,"distributed_executor":true}` | `{"cdc":true,"role":"elt","lineage":false,"paradigm":"both","dbt_native":true,"data_quality":false,"elt_pushdown":true,"python_first":false,"managed_cloud":true,"connector_count":"900+ sources and destinations","self_hosted_oss":false,"incremental_runs":true,"distributed_executor":true}` |
| integrations.count | 5 | 900 |
| integrations.list | `[{"name":"Hadoop"},{"name":"HDFS"},{"name":"YARN"},{"name":"Kubernetes"},{"name":"Docker"}]` | `[{"name":"Snowflake"},{"name":"Google BigQuery"},{"name":"Salesforce"},{"name":"Databricks"},{"name":"AWS Redshift"},{"name":"Azure Synapse"},{"name":"Marketo"},{"name":"HubSpot"},{"name":"Zendesk"},{"name":"Braze"},{"name":"Facebook Ads"},{"name":"Google Ads"},{"name":"Postgres"},{"name":"MySQL"}]` |
| language.primary | Scala | - |
| license.spdx | Apache-2.0 | - |
| market.availability | `{"primaryMarkets":[],"availabilityScope":"global","availableCountries":[],"notAvailableCountries":[]}` | `{"hqCountry":"US","primaryMarkets":["US","EU"],"availabilityScope":"global","availableCountries":[],"notAvailableCountries":[]}` |
| platform.support | `{"cli":true}` | `{"web":true}` |
| pricing | `{"type":"open_source","freeTier":true,"sourceUrl":"https://spark.apache.org","retrievedAt":"2026-08-14T13:48:09.654Z"}` | `{"type":"usage","plans":[{"free":true,"name":"Free","summary":"Free up to 500k MAR connections, 3.5k MAR activations, 5k MMR transformations","features":["Access to core platform","700+ connectors","200+ activation destinations","dbt Core integration"],"description":"Introductory offering for low-volume data with core platform functionality","contactSales":false,"includedLimits":{"activations_mar":"3.5k/month","connections_mar":"500k/month","transformations_mmr":"5k/month"}},{"free":false,"name":"Standard","summary":"Usage-based pricing on Monthly Active Rows; $5 base charge per connection (1-1M MAR range)","features":["Unlimited users","15-minute syncs","700+ fully managed connectors","200+ fully managed activation destinations","dbt Core integration","Role-based access control","REST API access","SSH tunnels"],"commitment":"monthly","components":[{"kind":"fixed","amount":5,"period":"month","currency":"USD"}],"description":"Core platform functionality for teams automating data movement","contactSales":false,"includedLimits":{}},{"free":false,"name":"Enterprise","summary":"All Standard features plus premium sync speed and deployment flexibility","features":["All Standard features","1-minute syncs","Fivetran Activations Audience Hub","Enterprise database connectors","Custom roles","VPN tunnels","SCIM/user provisioning","Multi-cloud provider choice (GCP/AWS/Azure)","Hybrid deployment"],"commitment":"annual","description":"Greater platform flexibility and granular control for scale","contactSales":false,"includedLimits":{}},{"free":false,"name":"Business Critical","summary":"Enterprise-grade security with customer-controlled encryption and advanced compliance","features":["All Enterprise features","Customer-managed keys for encryption","PCI DSS Level 1 certification","Private networking options"],"commitment":"annual","description":"Highest data protection and compliance for regulated or sensitive data","contactSales":false,"includedLimits":{}}],"addOns":[{"name":"Professional Services"},{"name":"Enterprise License Agreement"}],"summary":"Free tier available. Usage-based pricing starting $5/month base + per-MAR metering. Annual contracts save up to 22%. Enterprise License Agreements available for fixed annual pricing.","currency":"USD","freeTier":true,"sourceUrl":"https://www.getcensus.com/pricing","retrievedAt":"2026-08-13T22:21:48.361Z","freeTrialDays":14,"startingPrice":{"amount":5,"period":"month","currency":"USD"},"billingPeriods":["month","year"]}` |
| pricing.free_tier | yes | yes |
| pricing.model | commercial | freemium |
| pricing.price_level | free | low |
| pricing.transparent | yes | yes |
| reliability.status_page | - | yes |
| security.disclosure_policy | yes | - |
| security.gdpr | - | yes |
| security.hipaa | - | yes |
| security.iso27001 | - | yes |
| security.pci | - | yes |
| security.scorecard | 5.6 | - |
| security.soc2 | - | yes |
| security.vulnerabilities | `{"count":7,"source":"https://advisories.ecosyste.ms/api/v1/advisories?ecosystem=maven&package_name=org.apache.spark%3Aspark-core_2.10&per_page=100","last_12m":1,"max_severity":"CRITICAL"}` | - |

## Capabilities (Data Engineering Tools)

| Capability | Apache Spark | Census |
|---|:--:|:--:|
| **Core** |  |  |
| Tool role | Processing engine | ELT / ingestion |
| Processing paradigm | Batch + streaming | Batch + streaming |
| **Connectivity** |  |  |
| Connector count | - | 900+ sources and destinations |
| CDC / log-based replication | - | ✓ |
| **Transformation** |  |  |
| In-warehouse transformation (push-down) | - | ✓ |
| dbt-native orchestration | - | ✓ |
| **Deployment** |  |  |
| Self-hosted / open-source available | ✓ | ✗ |
| Managed cloud available | - | ✓ |
| **Governance** |  |  |
| Data lineage / asset catalog | - | ✗ |
| **Authoring** |  |  |
| Python-first authoring | - | ✗ |
| **Execution** |  |  |
| Incremental / partition-aware runs | - | ✓ |
| **Quality** |  |  |
| Built-in data quality / tests | - | ✗ |
| **Scale** |  |  |
| Horizontal scale (distributed executor) | ✓ | ✓ |

*Source: Vioscale. Generated 2026-09-01T14:39:18.905Z. "-" = undocumented, not absent.*
