Executive summary
Most analytics platforms were designed around one assumption: to analyze data, you first copy it into the vendor's cloud. For a large share of business data that assumption is fine. For the rest — customer records under strict residency rules, product databases with personal data, regulated financial ledgers, research datasets under contract — it creates a permanent tension between the people who want answers and the people accountable for where data goes.
This paper argues that the tension is an architecture problem with an architecture answer. A hybrid analytics architecture treats "where does this data live while we analyze it?" as a per-source setting rather than a platform-wide given. Sensitive sources stay on infrastructure you control and are queried live; their raw rows never come to rest on the vendor side. Less sensitive or high-volume event sources sync into a managed cloud, where history, caching and heavy joins are cheap.
Kimo implements the "stay home" half with Kimo Bridge, a small package you install next to your database. The bridge connects outward to Kimo over a single encrypted, mutually authenticated tunnel, so you open no inbound firewall ports. Kimo sends it queries generated from your semantic layer; the bridge checks each one against a policy that lives on your server, runs it with a read-only role, and returns only the result — usually an aggregate of a few hundred rows. Every query is logged locally, and you can revoke the bridge instantly.
The chapters that follow cover the regulatory context, the three architecture patterns and their trade-offs, the mechanics of the bridge, performance, deployment, a STRIDE threat model, a compliance mapping and a decision framework you can apply to your own source inventory. The paper ends with an honest section on limits: a bridge is not a substitute for a lawful basis, a data protection impact assessment or good database hygiene.
Why is data residency an architecture problem?
Data residency is the requirement — legal, contractual or self-imposed — that certain data be stored and processed in a particular place. For organizations subject to the EU General Data Protection Regulation (GDPR), the starting point is Chapter V. Article 44 sets the general principle: a transfer of personal data to a third country may take place only if the conditions of that chapter are met, including for onward transfers, so that the level of protection the Regulation guarantees is not undermined.1Source 1 · EUR-Lex, Official Journal of the European Union, 2016Regulation (EU) 2016/679 (General Data Protection Regulation)eur-lex.europa.eu
In practice that means one of three things. The destination has an adequacy decision from the European Commission; or the exporter puts in place "appropriate safeguards" under Article 46 — most commonly the standard contractual clauses adopted by the Commission — together with enforceable data subject rights and effective legal remedies; or a narrow derogation applies.1Source 1 · EUR-Lex, Official Journal of the European Union, 2016Regulation (EU) 2016/679 (General Data Protection Regulation)eur-lex.europa.eu
Schrems II turned paperwork into engineering
On 16 July 2020, in Case C-311/18 (commonly called Schrems II), the Court of Justice of the European Union invalidated the EU-US Privacy Shield adequacy decision, while holding that the Commission's decision on standard contractual clauses remained valid.2Source 2 · Court of Justice of the European Union, 2020Press Release No 91/20: Judgment in Case C-311/18, Data Protection Commissioner v Facebook Ireland and Maximillian Schremscuria.europa.eu The European Data Protection Board (EDPB) clarified what that means for exporters: before transferring under standard contractual clauses, exporter and importer must verify that the destination country offers a level of protection essentially equivalent to that guaranteed in the EU, add supplementary measures where needed, and suspend the transfer where that protection cannot be ensured.3Source 3 · European Data Protection Board, 2020Frequently Asked Questions on the judgment of the CJEU in Case C-311/18edpb.europa.eu
The EDPB's follow-up Recommendations 01/2020 turned this into a six-step roadmap: know your transfers, identify the transfer tools, assess whether the tool is effective in light of the destination's laws and practices, adopt supplementary measures, take the procedural steps, and re-evaluate at appropriate intervals.4Source 4 · European Data Protection Board, 2021Recommendations 01/2020 on measures that supplement transfer tools, version 2.0edpb.europa.eu Two details from that document matter for analytics architects:
- Remote access counts. The Recommendations note that remote access by an entity from a third country to data located in the EEA is also considered a transfer.4Source 4 · European Data Protection Board, 2021Recommendations 01/2020 on measures that supplement transfer tools, version 2.0edpb.europa.eu "We store it in Frankfurt" does not settle the question if a tool in another jurisdiction reads it.
- Encryption in transit and at rest is not always enough. For a cloud provider that needs data in the clear to do its job, the EDPB states that transport and at-rest encryption, even taken together, do not constitute a sufficient supplementary measure if the importer holds the keys, in the scenarios where problematic third-country law applies.4Source 4 · European Data Protection Board, 2021Recommendations 01/2020 on measures that supplement transfer tools, version 2.0edpb.europa.eu
The legal landscape keeps moving. The European Commission adopted a new adequacy decision for the EU-US Data Privacy Framework on 10 July 2023, which covers transfers to US companies participating in that framework.5Source 5 · European Commission, 2023EU-US data transferscommission.europa.eu That helps many teams — but the Court of Justice has invalidated an EU-US adequacy decision twice before — Safe Harbour in 2015 and Privacy Shield in 2020 — and the framework covers only participating organizations.2Source 2 · Court of Justice of the European Union, 2020Press Release No 91/20: Judgment in Case C-311/18, Data Protection Commissioner v Facebook Ireland and Maximillian Schremscuria.europa.eu Architects who want to stay robust across the next ruling design so that the most sensitive data does not depend on any single transfer mechanism.
Minimization is the quiet requirement
Article 25(2) of the GDPR requires data protection by default: only personal data necessary for each specific purpose should be processed, which applies to the amount collected, the extent of processing, the storage period and accessibility.1Source 1 · EUR-Lex, Official Journal of the European Union, 2016Regulation (EU) 2016/679 (General Data Protection Regulation)eur-lex.europa.eu Recital 26 adds that the principles of data protection do not apply to anonymous information — while pseudonymised data remains personal data.1Source 1 · EUR-Lex, Official Journal of the European Union, 2016Regulation (EU) 2016/679 (General Data Protection Regulation)eur-lex.europa.eu Analytics usually needs aggregates, not individuals. An architecture that moves aggregates instead of rows is therefore not just faster; it is closer to what the law asks for by default.
Residency is not only a European concern. Sector rules, customer contracts, public-sector procurement and simple internal policy ("production data never leaves the VPC") produce the same constraint. The architecture question is identical in every case: can we answer the business's questions without moving the data that is not allowed to move?
Three architectures for analytics: cloud ELT, federation and hybrid
Every analytics stack answers two questions: where does computation happen, and where does data rest between queries? Three patterns cover nearly every real deployment.
1. Full cloud ELT
Extract data from each source, load it into a cloud warehouse or the analytics vendor's storage, and transform it there (ELT). Incremental syncs — often driven by change data capture — keep it fresh. This is the dominant pattern because it is simple to reason about: one place, one SQL dialect, cheap joins across sources, full history, and queries that never touch production systems.
The cost is that every synced row now lives in a second location, under a second set of access controls, with a second retention policy to manage. For sensitive sources that is precisely the copy your data protection officer does not want.
2. Query federation with pushdown
Leave data where it is and send queries to it. A federated engine translates a logical query into source-specific SQL and pushes as much work as possible into the source — filters, column selection, aggregations, joins within one database, limits. The Trino documentation describes the payoff plainly: better query performance, less network traffic between the engine and the source, and less load on the source.14Source 14 · Trino DocumentationPushdowntrino.io See query pushdown for the mechanics.
Federation keeps data at rest where it already lives, but historically required the engine to reach into your network — an inbound connection, a VPN, an allow-listed IP — and put real load on operational databases.
3. Hybrid, decided per source
Hybrid means each source picks its mode. Your Postgres customer database is queried live; your marketing platforms, product events and billing exports sync to the cloud. The semantic layer hides the difference from users: "net revenue retention by region" resolves to the same measure whichever path the underlying tables take.
| Criterion | Full cloud ELT | Federation / pushdown | Hybrid (per source) |
|---|---|---|---|
| Query latency | Lowest for heavy joins and scans; data is co-located | Depends on source speed and network; fast when aggregates are pushed down | Cloud speed for synced sources, source speed for bridged ones |
| Freshness | As fresh as the sync schedule (minutes to hours) | Real time: every query reads current data | Chosen per source |
| Cost profile | Storage plus sync compute; grows with data volume | Compute on your database; grows with query volume | Pay storage only where history or speed earns it |
| Security surface | Second copy of all data under vendor controls | No second copy; exposure is query results | Second copy only for sources you choose to sync |
| Residency and compliance | Raw rows leave your environment | Raw rows stay; results move | Strictest sources stay; the rest move under existing safeguards |
| History and time travel | Native: snapshots and incremental history | Only what the source keeps | Native for synced sources; source-dependent for bridged |
| Cross-source joins | Cheap | Expensive: rows must meet somewhere | Cheap after aggregation; design measures accordingly |
| Load on production | Sync reads only | Every query | Bounded by per-source rate limits and replicas |
The table explains why most teams who think carefully end up hybrid. Pure ELT optimizes for analyst convenience and pays in exposure. Pure federation optimizes for exposure and pays in latency and production load. Hybrid lets you spend each budget only where it buys something.
How Kimo Bridge works
Kimo Bridge is a small service that holds no business data — its only state is its own identity — distributed as the container image ghcr.io/getkimo/bridge, a Helm chart and a single static binary called kimo-bridge — that you run inside your network, close to the databases it serves. It has three jobs: hold a secure connection to Kimo, enforce your policy on every request, and execute approved queries against your sources with credentials that never leave your server.
Scroll sideways to see the full diagram.
An outbound-only tunnel
The bridge initiates a single long-lived connection from your network to Kimo's regional bridge endpoint on TCP port 443 — the pattern described in our glossary entry on outbound-only tunnels. Because the connection originates inside your perimeter, you add no inbound firewall rules, no port forwards and no public IP. Your firewall sees one egress flow to one hostname, which you can allow-list and monitor like any other.
The connection uses TLS 1.3, the version of the protocol designed to prevent eavesdropping, tampering and message forgery between client and server.7Source 7 · IETF / RFC Editor, 2018RFC 8446: The Transport Layer Security (TLS) Protocol Version 1.3rfc-editor.org TLS 1.3 supports certificate-based client authentication: the server sends a CertificateRequest, and the client proves possession of its private key.7Source 7 · IETF / RFC Editor, 2018RFC 8446: The Transport Layer Security (TLS) Protocol Version 1.3rfc-editor.org Kimo Bridge always uses it — that is the "mutual" in mutual TLS (mTLS). The bridge generates its key pair locally at enrollment; the private key never leaves your host. Kimo issues a short-lived client certificate bound to that key and your workspace, and the bridge rotates it automatically.
This follows the core idea of zero trust architecture as NIST defines it: no implicit trust granted to assets or user accounts based solely on their physical or network location.6Source 6 · National Institute of Standards and Technology, 2020SP 800-207: Zero Trust Architecturecsrc.nist.gov Being "inside the tunnel" does not authorize anything. Each request is authenticated and authorized on its own.
The life of a query
- Step 1:
A question becomes a governed query
A user opens a dashboard or asks Ask Kimo a question. Kimo resolves every measure and dimension against the semantic layer and compiles SQL for the target source's dialect, pushing filters, projections and aggregations into the query.
- Step 2:
Kimo signs the request
The request carries the workspace, the requesting user, the source, a hash of the SQL and the policy version Kimo expects. Kimo signs it; the bridge verifies the signature against a key pinned at enrollment.
- Step 3:
The bridge enforces local policy
The bridge parses the SQL, rejects anything that is not a single read statement, checks tables and columns against your allow-list, injects row filters for the requesting user's attributes, and applies row, byte and time caps. Local policy always wins: Kimo can narrow it, never widen it.
- Step 4:
The database executes under a read-only role
The bridge runs the query with a dedicated read-only role, ideally against a replica, under a statement timeout. Database-native controls such as row-level security apply on top of the bridge's own.
- Step 5:
Only the result returns
The bridge streams the result back through the tunnel and writes an audit record locally: who, what, when, which policy, how many rows, how long. Kimo renders the chart. In Bridge mode the result is kept, at most, in a short-lived cache you can disable.
Two modes, chosen per source
Every source behind a bridge runs in one of two modes. Bridge mode is live query pushdown: nothing is copied, every query reads current data, and Kimo stores only short-lived result caches (configurable down to zero). Cloud mode uses the same outbound tunnel to sync tables incrementally into Kimo's managed cloud, which gives you history, faster heavy queries and less load on the source. The same bridge can serve both: your customers database in Bridge mode, your events warehouse in Cloud mode.
bridge:
name: fra-prod-01
region: eu
sources:
- id: crm_pg
type: postgres
dsn: ${CRM_READONLY_DSN}
mode: bridge # live pushdown, nothing stored by Kimo
cache_ttl: 0s # disable result caching entirely
- id: events_ch
type: clickhouse
dsn: ${EVENTS_READONLY_DSN}
mode: cloud # incremental sync to Kimo's managed cloud
sync:
schedule: '*/15 * * * *'
tables: [page_views, signups]Connector coverage behind the bridge mirrors Kimo's direct database connectors. The most common pairings are listed below; each links to its integration page.
Controls that travel with every query
Broken access control sits at the top of the OWASP Top 10 (2021), and OWASP's first prevention advice is simple: except for public resources, deny by default.8Source 8 · OWASP Top 10, 2021A01:2021 – Broken Access Controltop10.owasp.org Kimo Bridge is built around that rule. A freshly installed bridge can reach your database but answer nothing until you allow specific tables and columns.
| Control | Where it is enforced | What it stops |
|---|---|---|
| Read-only database role | Your database | Writes, DDL and privilege changes, even if every other layer failed |
| Single read statement only | Bridge SQL parser | Stacked statements, COPY, function calls with side effects |
| Table and column allow-list | Bridge policy | Reads of sensitive columns (e.g. email, IBAN) that no measure needs |
| Row filters | Bridge policy, plus database RLS | A regional manager seeing other regions' customers |
| Minimum group size | Bridge policy | Re-identification through very small aggregates |
| Result caps (rows, bytes, time) | Bridge and database timeout | Bulk export through the analytics path; runaway queries |
| Rate limits | Bridge | Load spikes on production; automated scraping |
| Signed requests and user attribution | Bridge | Unattributed or replayed queries |
| Local audit log | Your server | Unaccountable access; supports investigations and DPIAs |
| Kill switch | Your server and Kimo console | Continued access once you decide to stop it |
Read-only roles and database-native guardrails
The bridge connects with a role you create for it, and that role is the last line of defense. In PostgreSQL, default_transaction_read_only makes new transactions read-only by default and statement_timeout aborts any statement that runs too long;11Source 11 · PostgreSQL DocumentationClient Connection Defaultspostgresql.org ALTER ROLE … SET attaches both to the bridge role as session defaults applied at every login.16Source 16 · PostgreSQL DocumentationALTER ROLEpostgresql.org Session defaults can be overridden with SET, which is one reason the bridge parser rejects SET and set_config(). The predefined role pg_read_all_data grants read access broadly, but it does not bypass row-level security — useful to know when combining the two.10Source 10 · PostgreSQL DocumentationPredefined Rolespostgresql.org We still recommend granting SELECT on specific schemas instead: least privilege is easier to audit than a broad role you then fence in.
CREATE ROLE kimo_bridge LOGIN PASSWORD :'bridge_password';
ALTER ROLE kimo_bridge SET default_transaction_read_only = on;
ALTER ROLE kimo_bridge SET statement_timeout = '30s';
GRANT USAGE ON SCHEMA analytics TO kimo_bridge;
GRANT SELECT ON ALL TABLES IN SCHEMA analytics TO kimo_bridge;
ALTER DEFAULT PRIVILEGES IN SCHEMA analytics
GRANT SELECT ON TABLES TO kimo_bridge;Where tables hold rows that different people may see, PostgreSQL row-level security adds a second, independent filter. Once enabled on a table, a default-deny policy applies when no policy exists, meaning no rows are visible.9Source 9 · PostgreSQL DocumentationRow Security Policiespostgresql.org Superusers and roles with BYPASSRLS always bypass it, and table owners do too unless the table is set to FORCE ROW LEVEL SECURITY — so the bridge role must be neither.9Source 9 · PostgreSQL DocumentationRow Security Policiespostgresql.org
Policy that lives on your server
Bridge policy is a file on your host, versioned with the rest of your infrastructure code. Kimo shows the effective policy in the Bridge console but cannot change it: a change requires editing the file and reloading the bridge. That asymmetry is deliberate. If a Kimo account were compromised, the attacker could ask for less than your policy allows, never more.
policy:
default: deny
tables:
analytics.customers:
columns: [id, region, plan, mrr_cents, created_at, churned_at]
row_filter: "region = ANY({{user.attributes.regions}})"
analytics.invoices:
columns: [id, customer_id, amount_cents, currency, paid_at]
min_group_size: 5
limits:
max_rows: 50000
max_bytes: 20MB
timeout: 30s
rate: 20/sAudit and the kill switch
Every request — allowed or denied — produces a JSON audit record on your server with the requesting user, the source, the SQL hash, the policy version, the decision, row count, bytes and duration. OWASP recommends logging access control failures and alerting on repeated ones, and rate-limiting API access to reduce the harm from automated tooling;8Source 8 · OWASP Top 10, 2021A01:2021 – Broken Access Controltop10.owasp.org the bridge does both, and mirrors a copy of each record to Activity so analysts and auditors read the same trail.
Stopping access takes one action, from either side: run kimo-bridge pause on the host, revoke the bridge in the Kimo console (which revokes its certificate), or stop the container. The tunnel closes, in-flight queries are cancelled, and nothing in Bridge mode remains to be deleted beyond short-lived caches, which are purged on revocation.
How fast is live querying through a bridge?
Fast enough for interactive dashboards, when the work is pushed down. A typical dashboard tile — "monthly recurring revenue by plan for the last 24 months" — scans a large table but returns a few dozen rows. If the source computes the aggregate, the bridge moves kilobytes. If the engine fetches raw rows to aggregate them remotely, it moves gigabytes. Pushdown is the difference.
What the compiler pushes down
- Projection: only the columns a measure needs are selected — which also keeps disallowed columns out of the query entirely.
- Predicates: date ranges, segment filters and row filters run in the source's
WHEREclause, where indexes help. - Aggregation:
SUM,COUNT,COUNT(DISTINCT)and percentiles run at the source and return one row per group. - Joins within one source: a measure joining
invoicestocustomersin the same database is executed there. - Limits and top-N: "top 10 accounts by expansion" returns ten rows, not the whole ranking.
What cannot be pushed down is a join across two different sources. In hybrid deployments, design measures so that each side aggregates first and the two aggregates join in Kimo on a shared dimension (month, region, plan). That keeps both the latency and the data movement small.
Caching choices
| Cache setting | What is stored on Kimo's side | When to use it |
|---|---|---|
cache_ttl: 0s | Nothing beyond the response in flight | Most sensitive sources; strict residency |
cache_ttl: 5m (default) | Aggregated results for five minutes, encrypted | Shared dashboards opened by many people |
cache: local | Nothing; the bridge caches results on your host | Repeat queries where results must not leave the host at rest |
mode: cloud | Synced tables with history | Event data, marketing data, heavy exploratory work |
Protecting production
Point the bridge at a read replica whenever one exists. Then bound the blast radius twice: per-source rate limits in the bridge, and a role-level statement_timeout in the database.11Source 11 · PostgreSQL DocumentationClient Connection Defaultspostgresql.org For MySQL, account-level resource limits such as MAX_USER_CONNECTIONS and MAX_QUERIES_PER_HOUR play the same role.15Source 15 · MySQL 8.4 Reference ManualCREATE USER Statementdev.mysql.com Teams that follow these three habits rarely notice bridge load at all.
Deployment patterns
The bridge keeps no data of its own, so deployment is mostly a question of where it sits relative to your data and how many copies you run.
| Pattern | Shape | Good for |
|---|---|---|
| Single container | One docker run on a VM next to the database | Pilots, small teams, one database. See Install with Docker. |
| Highly available group | Two or more bridges with the same group name; Kimo balances queries across connected members | Production dashboards that must survive a host restart |
| Kubernetes | Helm chart, Deployment with 2+ replicas, PodDisruptionBudget, default-deny NetworkPolicy | Platform teams already running workloads on Kubernetes. See Deploy on Kubernetes. |
| One bridge per region | A bridge group in each region, each connected to Kimo's matching regional endpoint | Multinationals with data that must stay in-country |
| Segmented networks | A bridge in each network segment that holds data; no cross-segment routes required | Regulated environments with strict internal zoning |
On Kubernetes, the outbound-only property can be enforced by the cluster rather than trusted. Pods are non-isolated for egress until a NetworkPolicy with Egress in its policyTypes selects them; once selected, only the egress the policies list is allowed — but only if your network plugin actually enforces NetworkPolicy.12Source 12 · Kubernetes DocumentationNetwork Policieskubernetes.io A default-deny egress policy plus two narrow allow rules (DNS, Kimo's endpoint on 443) and one to the database turns the architecture diagram into something your cluster guarantees.
For environments that cannot make any outbound connection at all, a bridge is the wrong tool: run Kimo itself on your infrastructure instead, as described in our air-gapped deployment write-up and the on-premise documentation.
Production readiness checklist
- Bridge runs on a host or namespace dedicated to it, as a non-root user, with a read-only root filesystem.
- Egress is allowed only to Kimo's regional bridge hostname on TCP 443 (plus DNS).
- Database credentials come from a secret store, not from the command line or image.
- The database role is read-only, schema-scoped and has a statement timeout.
- The policy file is in version control and reviewed like code.
- At least two bridge instances in the same group for production sources.
- Audit logs ship to your SIEM; alerts fire on repeated denials.
- Someone owns the kill switch, and you have tested it.
Threat model: what could go wrong, and what stops it?
We use STRIDE, the threat model Microsoft uses in the Threat Modeling Tool of its Security Development Lifecycle: spoofing, tampering, repudiation, information disclosure, denial of service and elevation of privilege.13Source 13 · Microsoft LearnThreats — Microsoft Threat Modeling Tool (STRIDE model)learn.microsoft.com The table covers the bridge path specifically; it assumes the database itself is patched and its superuser credentials are protected.
| Threat (STRIDE) | Example on the bridge path | Primary mitigations | Residual risk |
|---|---|---|---|
| Spoofing | An attacker impersonates Kimo to the bridge, or a rogue host impersonates your bridge | mTLS with pinned server identity; client keys generated locally; per-request signatures | Compromise of the bridge host itself |
| Tampering | A query is altered in transit to read more data | TLS 1.3 integrity; SQL hash in the signed request; local policy re-checks the final SQL | Low |
| Repudiation | A user denies having run a sensitive query | User identity in each signed request; local, append-only audit log mirrored to Kimo | Depends on log retention you configure |
| Information disclosure | Bulk extraction through many "analytics" queries | Column allow-lists, row filters, minimum group size, row/byte caps, rate limits, alerting on denials | A permitted user can still see what policy permits |
| Denial of service | Expensive queries degrade the production database | Read replica, rate limits, statement timeout, connection caps | Replica lag under heavy load |
| Elevation of privilege | A query attempts writes, DDL or role changes | Single-statement read-only parser; read-only role; no superuser or BYPASSRLS | Database vulnerabilities outside the bridge's control |
The most important design property is that controls are layered and owned locally. The bridge's policy can fail open only if the database role also fails open; Kimo-side compromise can at worst issue queries your local policy already allows, each of which is logged on your server. The detailed control-by-control walkthrough lives in The Kimo Bridge security model.
Compliance mapping
An architecture does not make you compliant; it changes what you have to prove. The table maps common obligations to the evidence a hybrid deployment can produce. Kimo does not claim any certification on behalf of your deployment, and you should validate the mapping with your own counsel and auditors.
| Obligation | How the architecture helps | Evidence you can show |
|---|---|---|
| GDPR Art. 44–46: transfers need a legal basis and safeguards | In Bridge mode, raw rows stay in your environment; only policy-filtered results move | Source inventory with mode per source; policy files; data flow diagram |
| GDPR Art. 25(2): data protection by default | Deny-by-default policy; only columns needed by measures are selectable; aggregates by default | Column allow-lists; minimum group size settings |
| GDPR Art. 32: security of processing | Encryption in transit (TLS 1.3), least-privilege roles, access logging | Role definitions; audit log samples; key rotation records |
| EDPB Rec. 01/2020, Step 1: know your transfers | Every query and result is recorded with source, user and size | Audit log exports per period and per source |
| EDPB Rec. 01/2020, Step 6: re-evaluate periodically | Modes are a per-source setting you can switch without re-platforming | Policy history in version control |
| Internal policy: production data stays in the VPC | Outbound-only tunnel, no inbound rules, no replicated copy in Bridge mode | Firewall rules; NetworkPolicy manifests |
Article 32 lists the pseudonymisation and encryption of personal data among appropriate technical measures.1Source 1 · EUR-Lex, Official Journal of the European Union, 2016Regulation (EU) 2016/679 (General Data Protection Regulation)eur-lex.europa.eu A bridge policy can implement pseudonymisation at the source — for example by allowing only a hashed customer identifier instead of the email address — so that the analytics side never handles the direct identifier.
Should a source use Bridge mode or Cloud mode?
Decide per source, not per company. Score each source on five questions; the answers usually make the choice obvious.
Scroll sideways to see the full diagram.
| Question | Points toward Bridge mode | Points toward Cloud mode |
|---|---|---|
| 1. How sensitive is the data? | Direct identifiers, health, financial accounts, contractual restrictions | Aggregated platform metrics, public or low-sensitivity data |
| 2. Are there legal or contractual limits on location? | Yes, or the transfer assessment is uncertain | No, or a robust transfer mechanism is in place |
| 3. How fresh must answers be? | Real time matters (operations, finance close) | Fifteen-minute or hourly freshness is fine |
| 4. How heavy are the queries? | Aggregates over indexed tables; moderate volume | Wide scans, exploratory work, cross-source joins |
| 5. Do you need history the source does not keep? | No — the source already keeps it | Yes — snapshots, slowly changing dimensions, deleted rows |
Three outcomes are common. Bridge everything sensitive, sync everything else — the default hybrid. Bridge with a local cache for sensitive sources that power heavily shared dashboards. Cloud with a bridged dimension where event data syncs but the customer table that labels it stays home, with joins on non-identifying keys. If you are still unsure, start in Bridge mode: moving a source from Bridge to Cloud later is a one-line change, while un-copying data is much harder.
A worked example
Consider a fictional European B2B software company with four sources. Its Postgres product database holds customer contacts and usage; Stripe holds billing; HubSpot holds pipeline; a ClickHouse cluster holds 2 billion product events. Applying the framework: Postgres scores high on sensitivity and has contractual residency commitments, so it runs in Bridge mode with cache_ttl: 0s and an allow-list that excludes contact columns. ClickHouse holds pseudonymous events and serves heavy exploration, so it syncs in Cloud mode. Stripe and HubSpot connect directly as SaaS connectors. The board deck's net revenue retention combines a bridged aggregate (accounts by plan) with a synced one (billing), joined on month and plan.
Methodology and limits
Legal and regulatory statements in this paper come from primary texts: the GDPR on EUR-Lex, the Court of Justice press release on Case C-311/18, the EDPB's FAQ and Recommendations 01/2020, and the European Commission's page on EU-US transfers. Security statements are tied to standards and official documentation (NIST SP 800-207, the TLS 1.3 RFC, OWASP, PostgreSQL, MySQL and Kubernetes documentation). Performance figures are simulated and labelled as such; they illustrate the shape of the trade-off, not a benchmark you should plan capacity on.
- A bridge is not a legal basis. It reduces what moves and makes movement auditable. Whether a given flow is lawful remains a legal assessment.
- Results can be personal data. Policy design — column choices, group sizes — determines how much personal data appears in results.
- The bridge host is in scope. Whoever controls that host controls the credentials. Harden and monitor it like any production service.
- Live querying moves load to your database. Replicas, timeouts and rate limits keep it bounded, not zero.
- Regulation changes. The 2023 adequacy decision may not be the last word, which is one reason to keep the most sensitive sources independent of any single transfer mechanism.5Source 5 · European Commission, 2023EU-US data transferscommission.europa.eu
Conclusion
For a decade, analytics architecture has quietly assumed that the price of insight is a copy of your data in someone else's cloud. That price was acceptable for clickstreams and ad spend. It is much harder to justify for customer records, ledgers and regulated datasets — and since Schrems II, much harder to defend.
A hybrid architecture removes the false choice. Keep the data that must stay home on your infrastructure, push the computation to it, and move only the answers, under a policy you own and an audit trail you keep. Sync the rest for speed and history. Kimo Bridge is our implementation of that idea: one outbound connection, deny by default, read-only all the way down, and revocable in one action.
To see it with your own data, read the Kimo Bridge overview, follow the Docker install guide, or open the Bridge console in your workspace. For the shorter version of this argument, see Cloud, hybrid or bridge.
Sources (16)
Every factual claim above cites a numbered source. We link primary documents wherever they exist.
Sources
16 references- Regulation (EU) 2016/679 (General Data Protection Regulation) (opens in a new tab)EUR-Lex, Official Journal of the European Union2016eur-lex.europa.eu
Art. 44 general principle for transfers, Art. 46 appropriate safeguards, Art. 25(2) data protection by default, Art. 32 security of processing, Recital 26 anonymous information.
- Press Release No 91/20: Judgment in Case C-311/18, Data Protection Commissioner v Facebook Ireland and Maximillian Schrems (opens in a new tab)Court of Justice of the European Union2020curia.europa.eu
Privacy Shield decision invalidated; standard contractual clauses decision valid; recalls the 2015 Schrems I invalidation of Safe Harbour.
- Frequently Asked Questions on the judgment of the CJEU in Case C-311/18 (opens in a new tab)European Data Protection Board2020edpb.europa.eu
Exporters must verify essentially equivalent protection before transferring under SCCs, add supplementary measures, or suspend.
- Recommendations 01/2020 on measures that supplement transfer tools, version 2.0 (opens in a new tab)European Data Protection Board2021edpb.europa.eu
Six-step roadmap; remote access as a transfer; Use Case 6 on processors needing data in the clear.
- EU-US data transfers (opens in a new tab)European Commission2023commission.europa.eu
Adequacy decision for the EU-US Data Privacy Framework adopted 10 July 2023.
- SP 800-207: Zero Trust Architecture (opens in a new tab)National Institute of Standards and Technology2020csrc.nist.gov
No implicit trust based solely on physical or network location.
- RFC 8446: The Transport Layer Security (TLS) Protocol Version 1.3 (opens in a new tab)IETF / RFC Editor2018rfc-editor.org
TLS 1.3 goals (prevent eavesdropping, tampering, forgery) and certificate-based client authentication. Since obsoleted by RFC 9846, which keeps the same protocol version.
- A01:2021 – Broken Access Control (opens in a new tab)OWASP Top 102021top10.owasp.org
Ranked first in 2021; deny by default, log access control failures, rate-limit API access.
- Row Security Policies (opens in a new tab)PostgreSQL Documentationpostgresql.org
Default-deny when RLS is enabled without policies; superusers, BYPASSRLS and owners bypass unless FORCE.
- Predefined Roles (opens in a new tab)PostgreSQL Documentationpostgresql.org
pg_read_all_data does not bypass row-level security.
- Client Connection Defaults (opens in a new tab)PostgreSQL Documentationpostgresql.org
statement_timeout and default_transaction_read_only.
- Network Policies (opens in a new tab)Kubernetes Documentationkubernetes.io
Egress isolation and default-deny policies; requires a network plugin that enforces them.
- Threats — Microsoft Threat Modeling Tool (STRIDE model) (opens in a new tab)Microsoft Learnlearn.microsoft.com
Definitions of the six STRIDE categories.
- Pushdown (opens in a new tab)Trino Documentationtrino.io
Predicate, projection, aggregation, join, limit and top-N pushdown; benefits for performance, network traffic and source load.
- CREATE USER Statement (opens in a new tab)MySQL 8.4 Reference Manualdev.mysql.com
REQUIRE SSL/X509 and per-account resource limits.
- ALTER ROLE (opens in a new tab)PostgreSQL Documentationpostgresql.org
ALTER ROLE … SET stores a role-specific session default applied at login.
External sources were accessed at the time of writing. Kimo product details, customers and figures in examples are illustrative unless a source is cited.
Frequently asked questions
Does Kimo store my data when a source uses Bridge mode?
Do I need to open inbound firewall ports?
Does using a bridge make an international transfer lawful?
Will live queries slow down my production database?
Can Kimo change what my bridge is allowed to query?
Can I mix Bridge and Cloud modes?
Visser, A. (2026). Your Data, Your Rules. Kimo Research. https://getkimo.com/whitepapers/hybrid-analytics-architecture




