The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.

qio

CRAN status R-CMD-check coverage

qio reads and writes Apache Parquet files from R. Read a whole file with read_parquet(), or open a larger one with open_parquet() to inspect its schema and read only the columns, row groups, or batches you need. It is built on the bundled C library carquet and has no required R package dependencies.

Installation

Install the released version from CRAN:

install.packages("qio")

Or the development version from GitHub:

pak::pak("pedrobtz/qio")

Building from source needs GNU make and a C compiler. Zstandard and LZ4 are bundled; zlib comes from the system, or from Rtools on Windows.

Usage

library(qio)

write_parquet() and read_parquet() handle a whole file at a time.

write_parquet(mtcars, "mtcars.parquet")

cars <- read_parquet("mtcars.parquet")

Read part of a file by naming columns or row groups. A column you do not select is never decompressed, which is the cheapest speed-up available on a wide file.

read_parquet("mtcars.parquet", columns = c("mpg", "cyl"))

open_parquet() returns a handle. Inspecting one is cheap because it reads the footer, not the data, so you can look before deciding what to read.

# several row groups, so there is something to select
write_parquet(mtcars, "mtcars.parquet", row_group_size = 16)

pf <- open_parquet("mtcars.parquet")
pf
#> <qio_parquet_file>
#> mtcars.parquet
#> 32 rows x 11 columns; 2 row groups

schema(pf)      # column paths, Parquet types, nullability
row_groups(pf)  # rows and bytes per group
read_plan(pf)   # the R type each column will become

collect(pf, columns = c("mpg", "cyl"))
collect(pf, row_groups = 1)

close_parquet(pf)

Use walk_batches() for a file that does not fit in memory. It calls your function once per batch and keeps only one batch alive at a time.

pf <- open_parquet("big.parquet")

walk_batches(pf, batch_size = 100000, FUN = function(batch, index) {
  # one data frame at a time
})

close_parquet(pf)

Handles hold an open file, so close them when you are done. read_parquet() opens and closes one for you.

Type mapping

Parquet describes storage with a physical type and meaning with an optional logical type. qio reads them like this by default:

Parquet Logical type R
BOOLEAN logical
INT32 integer
INT32 DATE Date
INT32, INT64 TIME double (seconds)
INT64 double
INT64 TIMESTAMP POSIXct (UTC)
INT96 POSIXct (UTC)
FLOAT, DOUBLE double
BYTE_ARRAY STRING, ENUM, JSON character
BYTE_ARRAY list of raw
FIXED_LEN_BYTE_ARRAY list of raw
FIXED_LEN_BYTE_ARRAY UUID character
FIXED_LEN_BYTE_ARRAY FLOAT16 double
any DECIMAL double

Bytes are text only when the file says so, which is why an unannotated BYTE_ARRAY stays raw. INT64 is exact through 2^53 and NA beyond it; pass int64 = "integer64" for the full range. Nulls become NA. Nested and repeated columns are skipped for now.

Writing infers the reverse, and parquet_schema() overrides it per column:

R Parquet
logical BOOLEAN
integer INT32
double DOUBLE
character, factor BYTE_ARRAY + STRING
Date INT32 + DATE
POSIXct INT64 + TIMESTAMP (UTC, microseconds)
types <- parquet_schema(mpg = "FLOAT", cyl = "INT64")
write_parquet(mtcars, "mtcars.parquet", schema = types)

read_plan() reports what any file will produce before you read it, and ?qio-types documents every mapping, including where precision is lost.

Other functions

write_parquet() takes compression, row_group_size, metadata, sorted_by, and append. Beyond schema() and row_groups(), a handle can report column_chunks(), column_statistics(), page_index(), metadata(), and bloom_filter_may_contain(). validate_parquet() checks that a file is structurally sound without reading it.

See ?qio-limitations for what qio deliberately does not do.

Licensing

qio is MIT licensed. It bundles third-party C sources, each under its own license and shipped with its license file:

Bundled License Location
carquet MIT src/carquet/LICENSE
Snappy (in carquet) BSD-3-Clause src/carquet/compression/snappy.c
Zstandard BSD-3-Clause src/zstd/LICENSE
LZ4 BSD-2-Clause src/lz4/LICENSE

inst/COPYRIGHTS lists every copyright holder, the files each covers, and the modifications qio makes; DESCRIPTION points at it through its Copyright field. Each bundled library is pinned to an exact upstream commit.

qio carries local patches to carquet; most fix defects that silently corrupted or rejected valid data. They live as individual commits on the qio branch of a carquet fork, which is what src/carquet is vendored from, so each one can be read on its own and offered upstream. The reason for each is recorded in the repository’s .agents/VENDORED.md, which is not shipped in the source package.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.