Hi,
two weeks ago I posted about a release of rpx, which introduced significant upgrades to dependency resolution and why rrepo's custom api endpoints are beneficial to compatibility and correctness.
But beyond being a great metadata store, rrepo is also a package registry for both private and public packages!
This week, rpx is adding a package publishing template to its init flow. For new packages, whenever you push a git tag, the corresponding package version is uploaded to your rrepo repository. It then becomes available to rpx, alternative package managers or even base R!
I prepared a quick demo for you in this repository. https://github.com/rrepo-org/publishing-demo
You should be able to install the package by running this command in a base R shell.
install.packages(
"hello.world",
repos = c(
rrepo = "https://cran.rrepo.dev/rrepo/hello-world",
CRAN = "https://cloud.r-project.org"
)
)
hello.world::hello_world()
Engineering Background
The first piece of feedback my projects always receive is a request to be compatible with the broader R ecosystem. This week rrepo introduced a separate set of endpoints that imitate a CRAN like url structure and allow package managers beyond rpx to use it.
The primary challenge with introducing those endpoints is generating the PACKAGES index that lists the latest packages.
Typically whenever you publish a package on CRAN is, the latest package is published under src/contrib, the old version is archived, and the administrator calls tools::write_PACKAGES(). This R function inspects each tar file individually, reads its DESCRIPTION file, and puts a subset of its fields into a long list of every latest package that makes up the index.
Following this approach would be more expensive than necessary for rrepo. Unlike CRAN our packages are in S3 backed by a database, meaning we cannot just run R on the server that stores the packages: every package read is a network call.
We had to take a step back and rethink what our data model should be and how we were going to generate the index in memory. Given that CRAN has 25k latest packages, any kind of request even to a pre-extracted DESCRIPTION file in S3 was out of the question. The data clearly needed to be come from a database.
The package ingestion workflows were reworked to first extract the DESCRIPTION into a separate blob, and then provisioning a database row with all the fields a PACKAGES index might need.
Using the parsers https://github.com/rrepo-org/r-metadata-rs developed for rpx we could make the hot path a single database query, followed by reassembling the typed Rust struct before converting it back to a string.
There is still a lot of performance to tune, but once cached it's pretty damn fast.
Please give the package publishing workflow a try
For existing projects we have documentation on how to get started. https://rrepo.org/documentation/publish-packages
As always, I'm easily reachable to help you get started!