The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.
Starting with misha 5.3.0, databases can be stored in two formats:
The indexed format provides better performance and scalability, especially for genomes with many contigs (>50 chromosomes).
The indexed format uses unified files:
Sequence data: - seq/genome.seq - All
chromosome sequences concatenated - seq/genome.idx - Index
mapping chromosome names to positions
Track data: -
tracks/mytrack.track/track.dat - All chromosome data
concatenated - tracks/mytrack.track/track.idx - Index with
offset/length per chromosome
Advantages: - Fewer file descriptors (important for genomes with 100+ contigs) - Better performance for large workloads (14% faster) - Smaller disk footprint - Faster track creation and conversion
The per-chromosome format uses separate files:
Sequence data: - seq/chr1.seq,
seq/chr2.seq, … - One file per chromosome
Track data: -
tracks/mytrack.track/chr1.track, chr2.track, …
- One file per chromosome
When to use: - Compatibility with older misha versions (<5.3.0) - Small genomes (<25 chromosomes) where performance difference is negligible
By default, new databases use the indexed format:
Use gdb.info() to check your database format:
gdb.init_examples() # the bundled examples database
info <- gdb.info()
print(info$format) # "indexed" or "per-chromosome"
#> [1] "per-chromosome"The full record:
Convert all tracks and sequences to indexed format:
This will: 1. Convert sequence files (chr*.seq →
genome.seq + genome.idx) 2. Convert all tracks to indexed
format 3. Validate conversions 4. Remove old files after successful
conversion
Convert specific tracks while keeping others in legacy format:
Convert interval sets to indexed format:
# Only "big" interval sets are stored per chromosome, so lower the threshold
# to get big sets out of the small example database (default: 1,000,000).
options(gbig.intervals.size = 10)
gintervals.save("myintervals", gscreen("dense_track > 0.3"))
gintervals.save("my2dintervals", gextract("rects_track", gintervals.2d.all())[, 1:6])
options(gbig.intervals.size = 1e6)
# 1D intervals
gintervals.convert_to_indexed("myintervals")
# 2D intervals
gintervals.2d.convert_to_indexed("my2dintervals")High priority (significant benefits): - Genomes with many contigs (>50 chromosomes) - Large-scale analyses (10M+ bp regions frequently) - 2D track workflows - File descriptor limit issues
Medium priority (moderate benefits): - Repeated extraction workflows - Regular analyses on medium-sized regions (1-10M bp)
Low priority (minimal benefits): - Small genomes (<25 chromosomes) - One-off analyses - Simple queries on small regions
Step 1: Backup (optional but recommended)
# eval = FALSE: copies a whole database outside the vignette's temp directory.
# Create backup of important database
system("cp -r /path/to/mydb /path/to/mydb.backup")Step 2: Check current format
gdb.init_examples()
info <- gdb.info()
print(paste("Current format:", info$format))
#> [1] "Current format: per-chromosome"Step 3: Convert
Step 4: Verify
# Check format changed
info <- gdb.info()
print(paste("New format:", info$format))
#> [1] "New format: indexed"
# Test a few operations
result <- gextract("dense_track", gintervals(1, 0, 1000))
print(head(result))
#> chrom start end dense_track intervalID
#> 1 chr1 0 50 0.1777778 1
#> 2 chr1 50 100 0.1600000 1
#> 3 chr1 100 150 0.1800000 1
#> 4 chr1 150 200 0.1600000 1
#> 5 chr1 200 250 0.1600000 1
#> 6 chr1 250 300 0.2000000 1Step 5: Remove backup (after validation)
You can freely copy tracks between databases with different formats.
gtrack.copy()gtrack.copy(src, db = target_db) is the idiom. It reads
the source track directly, reconciles per-chromosome vs. indexed storage
and any difference in the two databases’ chromosome sets, and preserves
the track’s type. src may be a character vector, so a batch
copy is a single call.
gdb.init_examples()
source_db <- .misha$GROOT
# A second database to copy into
target_db <- file.path(tempdir(), "target_db")
unlink(target_db, recursive = TRUE)
gdb.create_linked(target_db, parent = source_db)
#> Created linked database at /dev/shm/aviezerl-rbuild/RtmpQxWCjh/target_db (linked to /dev/shm/aviezerl-rbuild/RtmpQxWCjh/trackdb/test)
# One track, or many
gtrack.copy("dense_track", db = target_db)
gtrack.copy(c("sparse_track", "array_track"), db = target_db)
gsetroot(target_db)
gtrack.ls()
#> [1] "array_track" "dense_track" "sparse_track"
gtrack.info("dense_track")[c("type", "size.in.bytes")]
#> $type
#> [1] "dense"
#>
#> $size.in.bytes
#> [1] 80012The destination does not have to be in the same format: converting it afterwards leaves the tracks readable.
gdb.convert_to_indexed()
gdb.info()$format
#> [1] "indexed"
head(gextract("dense_track", gintervals(1, 0, 300)))
#> chrom start end dense_track intervalID
#> 1 chr1 0 50 0.1777778 1
#> 2 chr1 50 100 0.1600000 1
#> 3 chr1 100 150 0.1800000 1
#> 4 chr1 150 200 0.1600000 1
#> 5 chr1 200 250 0.1600000 1
#> 6 chr1 250 300 0.2000000 1Exporting with gextract(..., file = ...) and
re-importing with gtrack.import() also works, but it is
lossy: binsize = 0 builds a sparse track,
so a dense source track comes back sparse and about 5x larger (20 bytes
per value against 4 bytes per bin). Reach for it only when the
destination is not a misha database, or when you actually want a
different track type.
gsetroot(source_db)
exported <- file.path(tempdir(), "dense_track.txt")
gextract("dense_track", gintervals.all(), iterator = "dense_track", file = exported)
gsetroot(target_db)
gtrack.import("dense_track_roundtrip", "Round-tripped", exported, binsize = 0)
gtrack.info("dense_track")[c("type", "size.in.bytes")]
#> $type
#> [1] "dense"
#>
#> $size.in.bytes
#> [1] 80012
gtrack.info("dense_track_roundtrip")[c("type", "size.in.bytes")]
#> $type
#> [1] "sparse"
#>
#> $size.in.bytes
#> [1] 400120
gtrack.rm("dense_track_roundtrip", force = TRUE)Based on comprehensive benchmarks comparing indexed vs legacy formats:
# Work with both formats in same session
gsetroot(source_db) # per-chromosome
data1 <- gextract("dense_track", gintervals(1, 0, 1000))
head(data1)
#> chrom start end dense_track intervalID
#> 1 chr1 0 50 0.1777778 1
#> 2 chr1 50 100 0.1600000 1
#> 3 chr1 100 150 0.1800000 1
#> 4 chr1 150 200 0.1600000 1
#> 5 chr1 200 250 0.1600000 1
#> 6 chr1 250 300 0.2000000 1
gsetroot(target_db) # indexed
data2 <- gextract("dense_track", gintervals(1, 0, 1000))
head(data2)
#> chrom start end dense_track intervalID
#> 1 chr1 0 50 0.1777778 1
#> 2 chr1 50 100 0.1600000 1
#> 3 chr1 100 150 0.1800000 1
#> 4 chr1 150 200 0.1600000 1
#> 5 chr1 200 250 0.1600000 1
#> 6 chr1 250 300 0.2000000 1This occurs with many-contig genomes in legacy format:
Solution: Convert to indexed format
After manually copying track directories:
Solution: Reload database
gdb.create_genome() for standard genomesgdb.create() with multi-FASTA for custom
genomesgdb.info()gdb.convert_to_indexed()These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.