The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.
qio reads and writes Apache Parquet files from R.
Read a whole file with read_parquet(), or open a larger one
with open_parquet() to inspect its schema and read only the
columns, row groups, or batches you need. It is built on the bundled C
library carquet and
has no required R package dependencies.
Install the released version from CRAN:
install.packages("qio")Or the development version from GitHub:
pak::pak("pedrobtz/qio")Building from source needs GNU make and a C compiler. Zstandard and LZ4 are bundled; zlib comes from the system, or from Rtools on Windows.
library(qio)write_parquet() and read_parquet() handle a
whole file at a time.
write_parquet(mtcars, "mtcars.parquet")
cars <- read_parquet("mtcars.parquet")Read part of a file by naming columns or row groups. A column you do not select is never decompressed, which is the cheapest speed-up available on a wide file.
read_parquet("mtcars.parquet", columns = c("mpg", "cyl"))open_parquet() returns a handle. Inspecting one is cheap
because it reads the footer, not the data, so you can look before
deciding what to read.
# several row groups, so there is something to select
write_parquet(mtcars, "mtcars.parquet", row_group_size = 16)
pf <- open_parquet("mtcars.parquet")
pf
#> <qio_parquet_file>
#> mtcars.parquet
#> 32 rows x 11 columns; 2 row groups
schema(pf) # column paths, Parquet types, nullability
row_groups(pf) # rows and bytes per group
read_plan(pf) # the R type each column will become
collect(pf, columns = c("mpg", "cyl"))
collect(pf, row_groups = 1)
close_parquet(pf)Use walk_batches() for a file that does not fit in
memory. It calls your function once per batch and keeps only one batch
alive at a time.
pf <- open_parquet("big.parquet")
walk_batches(pf, batch_size = 100000, FUN = function(batch, index) {
# one data frame at a time
})
close_parquet(pf)Handles hold an open file, so close them when you are done.
read_parquet() opens and closes one for you.
Parquet describes storage with a physical type and meaning with an optional logical type. qio reads them like this by default:
| Parquet | Logical type | R |
|---|---|---|
BOOLEAN |
logical |
|
INT32 |
integer |
|
INT32 |
DATE |
Date |
INT32, INT64 |
TIME |
double (seconds) |
INT64 |
double |
|
INT64 |
TIMESTAMP |
POSIXct (UTC) |
INT96 |
POSIXct (UTC) |
|
FLOAT, DOUBLE |
double |
|
BYTE_ARRAY |
STRING, ENUM, JSON |
character |
BYTE_ARRAY |
list of raw |
|
FIXED_LEN_BYTE_ARRAY |
list of raw |
|
FIXED_LEN_BYTE_ARRAY |
UUID |
character |
FIXED_LEN_BYTE_ARRAY |
FLOAT16 |
double |
| any | DECIMAL |
double |
Bytes are text only when the file says so, which is why an
unannotated BYTE_ARRAY stays raw. INT64 is
exact through 2^53 and NA beyond it; pass
int64 = "integer64" for the full range. Nulls become
NA. Nested and repeated columns are skipped for now.
Writing infers the reverse, and parquet_schema()
overrides it per column:
| R | Parquet |
|---|---|
logical |
BOOLEAN |
integer |
INT32 |
double |
DOUBLE |
character, factor |
BYTE_ARRAY + STRING |
Date |
INT32 + DATE |
POSIXct |
INT64 + TIMESTAMP (UTC, microseconds) |
types <- parquet_schema(mpg = "FLOAT", cyl = "INT64")
write_parquet(mtcars, "mtcars.parquet", schema = types)read_plan() reports what any file will produce before
you read it, and ?qio-types documents every mapping,
including where precision is lost.
write_parquet() takes compression,
row_group_size, metadata,
sorted_by, and append. Beyond
schema() and row_groups(), a handle can report
column_chunks(), column_statistics(),
page_index(), metadata(), and
bloom_filter_may_contain(). validate_parquet()
checks that a file is structurally sound without reading it.
See ?qio-limitations for what qio deliberately does not
do.
qio is MIT licensed. It bundles third-party C sources, each under its own license and shipped with its license file:
| Bundled | License | Location |
|---|---|---|
| carquet | MIT | src/carquet/LICENSE |
| Snappy (in carquet) | BSD-3-Clause | src/carquet/compression/snappy.c |
| Zstandard | BSD-3-Clause | src/zstd/LICENSE |
| LZ4 | BSD-2-Clause | src/lz4/LICENSE |
inst/COPYRIGHTS lists every copyright holder, the files
each covers, and the modifications qio makes; DESCRIPTION
points at it through its Copyright field. Each bundled
library is pinned to an exact upstream commit.
qio carries local patches to carquet; most fix defects that silently
corrupted or rejected valid data. They live as individual commits on the
qio branch of a carquet
fork, which is what src/carquet is vendored from, so
each one can be read on its own and offered upstream. The reason for
each is recorded in the repository’s .agents/VENDORED.md,
which is not shipped in the source package.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.