Single-cell reconstruction: methods and parameters¶
Two methods compute 3-D structures of one cell from its contacts, both on the native engine
uchrom-recon (Rust; multi-core CPU in f64, GPU in f32 through wgpu: Metal on macOS, Vulkan on Linux),
both returning an ensemble of models as one ChromData.
Method |
Module |
Best for |
Notes |
|---|---|---|---|
EMber |
|
single-cell Hi-C / Dip-C, noisy contacts, phased diploid |
protocol |
NucDynamics |
|
single-cell Hi-C / Dip-C |
the 2017 method of Stevens et al.; annealed ensembles |
NucDynamics (single-cell)¶
A re-implementation of NucDynamics (Stevens et al. 2017, Nature 544:59,
doi:10.1038/nature21429; original code tjs23/nuc_dynamics, LGPL-3.0) in the
separate native package uchrom-recon (packages/uchrom-recon, Rust + PyO3):
pip install u-chrom[recon] # or: pip install ./packages/uchrom-recon
It computes an ensemble of structures at once — models in parallel on the CPU (f64), all models batched on the GPU (f32, wgpu: Metal on macOS, Vulkan on Linux/NVIDIA).
Input: .pairs / .pairs.gz, NCC (.ncc), a chr_A pos_A chr_B pos_B
table (the Stevens 2017 GEO files) or a DataFrame. Output: a ChromData
with one spot per particle and one trace per chromosome; coords = model 0
and layers["model_<k>"] = every model (units: particle radii).
import uchrom.datasets as ds
from uchrom.recon.sc.nucdyn import reconstruct_nucdyn
pairs = ds.fetch("stevens2017") / "GSM2219497_Cell_1_contact_pairs.txt.gz" # Stevens 2017 Cell 1 (GEO)
cd = reconstruct_nucdyn(pairs, n_models=10, device="auto", seed=1)
cd.write("cell1.chromdata.zarr")
CLI:
python -m uchrom.recon.sc.nucdyn contacts.pairs.gz out.chromdata.zarr \
--n_models=10 --device=gpu --seed=1
Key parameters:
Param |
Default |
Meaning |
|---|---|---|
|
10 |
Structures in the ensemble |
|
|
|
|
0 (all) |
CPU threads |
|
0 |
Random seed (CPU results are bitwise reproducible per seed) |
|
|
The 2017 code behind the published structures; |
|
8, 4, 2, 0.4, 0.2, 0.1 Mb |
Hierarchical schedule |
|
500 / 100 |
Annealing temperature steps per stage / MD steps per temperature |
|
5000 / 10 |
Annealing temperatures |
|
None |
Restrict the contacts, e.g. |
EMber (single-cell, noise-aware)¶
EMber (EM + annealing / sampling) runs on the same native engine as
NucDynamics (engine protocol "robust", alias "ember"). It keeps the
NucDynamics 2017 base and adds a mixture observation model — each contact is a
true proximity or a false contact; contact restraints are re-weighted by
their posterior probability of being true (EM; false-contact rate and contact
kernel learned from the data, plus local-support evidence) — a short schedule,
a finite-temperature sampling stage (calibrated ensembles) and adaptive
resolution below 25 kb. It was pre-registered, frozen and tested against
NucDynamics on held-out contacts and imaging-truth simulations
(benchmarks/screcon/PREREG.md, TEST_RESULTS.md): better on both primary
endpoints, ~5x faster on one GPU; its real-data ensembles are less tight.
from uchrom.recon.sc import reconstruct_ember
cd = reconstruct_ember(pairs, n_models=10, device="auto", seed=1)
cd.results["ember.contact_weights"] # posterior weight of every input contact
# also: uchrom.tl.reconstruct_sc(contacts, method="ember")
# reconstruct_nucdyn(contacts, protocol="ember") (same coordinates for the same seed)
Programmatic API¶
from uchrom.recon.sc.nucdyn import reconstruct_nucdyn
cd = reconstruct_nucdyn("contacts.pairs.gz", n_models=10) # ChromData ensemble
cd.write("out.chromdata.zarr")
from uchrom import ChromData
cd = ChromData.read("out.chromdata.zarr")
Both write a ChromData with one trace per chromosome (coords = model 0, layers["model_<k>"] = every
model); downstream tools do not care which method made it. Bulk Hi-C: bulk Hi-C.