Databanking is a programme with several open research strands. Five are listed here, each with its own research object and each looking for a different kind of collaborator. The full project outlines are password-protected.
Sandbox security, audit evidence and multi-owner data.
The custody wall is not a feature of the architecture. It is the architecture, and it has to hold against an operator as well as against a requester, prove after the fact what it did, and cope with data that belongs to more than one person at once.
Open questions include the trust model for sandboxed execution, what an audit ledger must record to carry evidential weight, governance of jointly-owned records, and the conditions under which a model may be trained inside a Databank.
From legal norms to verifiable policy.
How can legal and regulatory requirements be transformed into formally verifiable executable policy while keeping the software implementation distinct from the law itself? The strand follows the path from authoritative legal source, through interpretation and formal specification, to runtime enforcement and audit evidence, treating ambiguity, discretion, provenance, certification and legal change as first-class problems.
Databanking is the initial application environment, but the underlying question is broader: any system that enforces a regulatory constraint has to decide what it does when the law runs out.
Placement, orchestration and the output channel.
Bringing the algorithm to the data leaves a further question open: which data, on whose hardware, and at what cost. Databanks may well be heterogeneous, and a single authorised query may span several of them.
This strand covers placement and orchestration of computation across heterogeneous infrastructure, and the communication problem at the boundary, where the result has to carry the meaning the requester needs while remaining bit-scarce.
Externally anchored, tamper-evident audit history.
A Databank's audit trail is only as trustworthy as the operator that keeps it. This strand studies periodic cryptographic commitments to the audit history, anchored outside the custodian's administrative control, so that retrospective alteration of committed records becomes independently detectable — while the audit records themselves, and the personal data behind them, never leave the Databank.
Open questions include commitment schemes for structured, low-entropy records, coupling log creation to output release so that omission becomes structurally difficult, anchoring cadence and metadata leakage, the choice among blockchains, timestamping authorities and transparency logs, and the evidential weight a valid proof should carry.
Energy and carbon cost of the custody model.
A copy of personal data is never merely kept. Every holder runs its own ingestion, cleaning, backup, monitoring and deletion machinery over substantially the same records. Databanking removes that duplication of processing and adds costs of its own, in sandboxed execution repeated per query. This strand builds the comparison and measures it.
Open questions include how to specify architectures that are genuinely comparable, what share of the removed cost is processing rather than storage, what sandboxing and attestation cost per query at realistic volumes, and at what query volume the architecture stops paying for itself.
The full outlines are access-controlled while the collaborations are being formed, as is the draft technical architecture paper that strands 1, 3 and 4 feed into. To read the full outlines, or to propose a strand that is not listed here, write to contact@databanking.org.