Migrate already-pinned IPFS CIDs onto Filecoin Onchain Cloud (FOC) without re-chunking.
Each original CID stays byte-for-byte intact and individually retrievable over IPFS, while far fewer pieces are committed on-chain. The storage provider pulls each object's bytes directly from a trustless IPFS gateway; your machine streams each object once to compute its piece commitment and stores none of the payload.
To run a passthrough migration with nothing installed — prepare and submit —
use the browser console —
sgtpooki.github.io/ipfs2foc. When a
run outgrows the tab, ipfs2foc serve runs the same console against a
local daemon; headless automation uses the
CLI commands directly.
- Install
- Requirements
- Prerequisites
- Quickstart
- Commands
- Recovery commands
- Troubleshooting
- How it works
- Aggregate lifecycle and park/commit safety
- Public ingress for provider pulls
- Network gas and payments
- State
- Scope and limits
- Documentation
- Roadmap
- Contributing
- License
npm install -g ipfs2foc # the `ipfs2foc` command
# or run without installing:
npx ipfs2foc --helpFrom source (development uses pnpm):
git clone https://github.com/SgtPooki/ipfs2foc
cd ipfs2foc
pnpm install
node packages/cli/src/index.ts --help # run directly; Node 26 strips the TypeScript types- Node 26+ (uses the built-in
node:sqlite). - A source that serves deterministic trustless CARs. The default is
trustless-gateway.link; others (for examplegateway.pinata.cloud) work via--gateway. A gateway that returns reassembled files instead of CARs does not work;probereports which case a gateway falls into. Seedocs/sources.mdfor per-provider notes and probe commands.
Before running the quickstart, complete the one-time wallet setup on the network
you target (default mainnet; pass --network calibration for the testnet).
-
Wallet: a wallet whose address is the FWSS payer for the data set you create or own. Export the key as
PRIVATE_KEY(0x+ 64 hex) in the environment. The same key signs the data-set creation, the pull authorization, and every AddPieces submission. -
FIL in that wallet for the migrator's own transactions: USDFC ERC-20 approve, FilecoinPay deposit, FilecoinWarmStorageService operator approval. These three steps happen once per payer; the storage provider pays gas for everything it submits on chain (createDataSet, AddPieces, proof of possession).
-
Payment setup: deposit USDFC into Filecoin Pay and approve FWSS as a payments operator with enough rate and lockup allowance, plus the minimum lockup and one-time sybil fee.
create-data-setreverts without it.filecoin-pindoes the deposit and approvals in one command:export PRIVATE_KEY=0x... npx filecoin-pin@latest payments setup --auto # add --network calibration for the testnet npx filecoin-pin@latest payments status # confirm the approvals and balance
The Synapse SDK
Paymentshelper exposes the same calls directly, and PDP Scan (https://pdp.vxb.ai/{network}) shows the resulting account state. -
Provider id: choose a PDP-capable provider from the SP registry. PDP Scan lists registered providers at
https://pdp.vxb.ai/{network}/providers; pass that numeric id as--provider-idtocreate-data-set. Skip this step if you are reusing an existing data set you own. -
Trustless gateway: confirm the gateway you intend to use returns byte-stable CARs for one of your CIDs with
probebefore runningplan.
Complete Prerequisites above first. Default network is mainnet; pass
--network calibration for the testnet. redirect-serve and pdp-submit run
concurrently in separate terminals.
First time? Rehearse the whole flow on the testnet with the calibration tutorial before spending real funds on mainnet.
export PRIVATE_KEY=0x...
# One-time payer setup: deposit USDFC and approve FWSS as a payments operator.
npx filecoin-pin@latest payments setup --auto
# 0. (Once per provider) Provision a data set with withIPFSIndexing. Note the
# printed `dataSetId`; reuse it in steps 4 and 5.
ipfs2foc create-data-set --provider-id <provider-id>
# 1. Confirm a trustless gateway returns a deterministic CAR for one of your CIDs.
ipfs2foc probe <sample-cid> --gateway https://trustless-gateway.link
# 2. Compute piece commitments and pack aggregates (one source CID per sub-piece).
printf '%s\n' <cid> > cids.txt
ipfs2foc plan --cids cids.txt --db migrate.db
# 3. (Terminal A — leave running) Serve sub-pieces over public HTTPS.
# `--ingress cloudflared` spawns a no-signup Cloudflare tunnel and logs the URL.
# (`ipfs2foc serve --ingress cloudflared` does the same from the console daemon.)
ipfs2foc redirect-serve --db migrate.db --port 4322 --ingress cloudflared
# 4. (Terminal B) Pull, park, and add each aggregate onto the provider's data set.
# `--source-base` is the public HTTPS origin only (no path) from step 3.
ipfs2foc pdp-submit --db migrate.db --data-set-id <data-set-id> \
--source-base https://<public-host>
# 5. Confirm every CID landed: reconcile local state against the on-chain pieces.
ipfs2foc report --db migrate.db --data-set-id <data-set-id>cids.txt: one CID per line; blank lines and # comments are ignored.
plan is INSERT-only: re-running it after appending CIDs adds new
sub-pieces and aggregates without rewriting prior planning state. Existing
submitted/parked/committed aggregates are never touched.
The quickstart runs the single-asset path: each source CID becomes one
passthrough sub-piece pulled straight from the gateway, with no staging disk.
Use the multi-asset path when source CIDs are smaller than the provider's
Min Piece Size or you want fewer on-chain pieces per source CID. It replaces
step 2 with a plan that defers packing, then assembles multi-root CARs on disk:
ipfs2foc plan --cids cids.txt --db migrate.db --no-auto-pack
ipfs2foc pack-cars --db migrate.db --car-store /var/foc-cars --pack-target-size 512MiBdocs/personas.md maps disk, bandwidth, and time budgets to
concrete knob settings for both paths.
# Check whether a gateway serves deterministic CARs for a CID
ipfs2foc probe <cid> [--gateway https://gateway.pinata.cloud]...
# Compute one PieceCID v2
ipfs2foc commp <cid> [--gateway URL]...
# Full pipeline: commitments + aggregate packing into a SQLite DB.
# Default auto-wraps each source CID as a passthrough sub-piece. Pass
# --no-auto-pack to defer sub-piece assembly to `pack-cars` (multi-asset).
ipfs2foc plan --cids cids.txt [--db migrate.db] [--gateway URL]... \
[--piece-size 32GiB] [--concurrency 8] [--no-auto-pack]
# Load a run manifest saved by the browser console: records its piece
# commitments as done pieces (recomputing nothing) and packs aggregates,
# leaving the DB as if `plan` had produced it. Refuses on network mismatch
# or a conflicting prior commitment; re-import is a no-op.
ipfs2foc import-manifest <manifest.json> [--db migrate.db] \
[--network mainnet|calibration] [--piece-size 32GiB] [--no-auto-pack]
# Multi-asset packer: assemble many source CIDs into one multi-root CAR per
# sub-piece, append aggregates over the new sub-pieces.
ipfs2foc pack-cars --db migrate.db --car-store <dir> [--gateway URL]... \
[--pack-target-size 512MiB] [--fetch-concurrency 4]
# Progress and the aggregate plan
ipfs2foc status [--db migrate.db] [--json]
# Pre-flight a CID list against a gateway: pass rate, sizes, throughput estimate
ipfs2foc analyze [--cids cids.txt] [--db migrate.db] [--car-store <dir>] [--gateway URL] \
[--sample 100|--all] [--probe-concurrency 8] [--bw-target URL] \
[--network mainnet|calibration] [--json]
# Background daemon + browser console (start/pause/resume, add CIDs, add gateways)
ipfs2foc serve [--db migrate.db] [--cids cids.txt] [--gateway URL]... \
[--port 4321] [--network mainnet|calibration] [--rpc-url URL] [--max-base-fee N] \
[--app-dir <dir>] [--ingress cloudflared | --public-base https://<host>]
# Current network base fee and whether to pause submission
ipfs2foc gas [--network mainnet|calibration] [--rpc-url URL] [--max-base-fee 1000000]
# Sub-piece server: GET /piece/{pcidv2} -> 302 to the gateway CAR for a
# passthrough sub-piece, or byte-serves the assembled CAR file for a
# multi-asset sub-piece.
ipfs2foc redirect-serve [--db migrate.db] [--port 4322] [--ingress funnel|cloudflared]
# Provision a new FWSS data set with withIPFSIndexing (PRIVATE_KEY env)
ipfs2foc create-data-set --provider-id <id> \
[--network mainnet|calibration] [--rpc-url URL] [--cdn] [--timeout-seconds 600]
# Migrate via the PDP pull path (PRIVATE_KEY env). The pull source is either
# your own redirect-serve origin (--source-base) or a shared stateless relay
# (--source-relay, passthrough sub-pieces only — no server of your own needed).
ipfs2foc pdp-submit --db migrate.db --data-set-id <id> \
(--source-base https://<public-host> | --source-relay https://<relay-base>) \
[--network mainnet|calibration] [--rpc-url URL] \
[--max-in-flight 4] [--max-base-fee 1000000] [--pull-batch 32] [--poll-seconds 15]
# Verification report: reconcile a run against the data set's on-chain pieces
ipfs2foc report --db migrate.db --data-set-id <id> \
[--network mainnet|calibration] [--rpc-url URL] [--json] \
[--check-ipni <delegated-routing-url>] [--ipni-sample 100|--ipni-all] [--ipni-concurrency 8]plan, commp, and serve also accept an opt-in IPFS fallback that recovers
from source-gateway 5xx/429 through an embedded node:
[--ipfs-fallback] [--ipfs-fallback-mode gateway-first] [--ipfs-fallback-timeout-seconds 120]pdp-submit honors the in-flight cap, the base-fee gate, and provider pull
backpressure (HTTP 429 + Retry-After). If the provider's add errors after the
on-chain AddPieces already landed, pdp-submit confirms the aggregate against the
data set's active pieces and marks it committed instead of adding it again.
A run prepared in the browser console submits the
same way: import-manifest the saved manifest, then pdp-submit --source-relay — the provider pulls each piece through the relay, so no
redirect server of your own is needed.
serve starts an HTTP server (default http://localhost:4321) that runs the commP pass
in the background and serves the bundled browser console as
its control plane: live piece counts, the aggregate plan with each aggregate's status
and parent CID, parked-but-uncommitted count, and failures. Controls: start, pause,
resume, retry failed, add CIDs (POST /api/cids), set gateways (POST /api/gateways).
All state lives in the DB, so the process can stop and resume.
The console asks GET /api/capabilities on load to discover what the backend can do;
the hosted copy of the same app gets a 404 there and falls back to its in-browser
prepare + signing flow. The server binds loopback only and rejects cross-origin
mutations. During development, point serve at a freshly built console with
--app-dir ../../app/dist or the IPFS2FOC_APP_DIR environment variable (build it
with base /, the way pnpm -C packages/cli build does — a GitHub Pages build uses
a different base path and will not load).
The console can also submit on chain without PRIVATE_KEY: connect a wallet in the
Signing panel and grant a session key (one wallet transaction, scoped to
CreateDataSet + AddPieces with an explicit expiry). The daemon verifies the grant on
chain, keeps the key in the migration database, and drives the provider pull/add
itself — the tab can close mid-run, and extending the session in the browser keeps a
long run going without restarting it. The stored key signs nothing beyond those two
operations and can be revoked from the console at any time; treat the .db file
like the working state it is.
These re-arm aggregates that did not reach committed. They are not part of a
routine migration; reach for them only when a run is stuck and you have read
docs/personas.md failure modes.
# Move `failed` aggregates back to `planned` so the next pdp-submit retries them
ipfs2foc reset-failed-aggregates [--db migrate.db] [--network mainnet|calibration]
# Re-arm `submitted`/`parked` aggregates that never confirmed.
# Only after verifying on chain that their roots are NOT present — re-arming an
# aggregate whose AddPieces actually landed lands a duplicate.
ipfs2foc retry-unconfirmed-aggregates [--db migrate.db] [--network mainnet|calibration]probereportsWARN. The gateway answered but the bytes do not re-hash to the requested CID, or the response is not a CAR. That gateway cannot be a source. Fix the gateway config (Kubo: setGateway.DeserializedResponsestofalse) or pick another fromdocs/sources.md.set PRIVATE_KEY (0x + 64 hex).create-data-setandpdp-submitread the signing key from the environment. Export it (export PRIVATE_KEY=0x...) orsource .envbefore running.create-data-setreverts. The payer's USDFC deposit, FWSS operator approval, or allowances are insufficient. See Prerequisites.- Provider rejects the pull / public-host error.
--source-basemust be the public HTTPS origin only (scheme + host, no path) and resolve to a public IP. CGNAT and private ranges are rejected. Seedocs/ingress.md. planreports a CID asoversized. Its padded piece size exceeds the--piece-sizeaggregate budget. A CAR above the provider's per-piece pull limit (~1 GiB raw) cannot be migrated as one piece either; hold it out of the run until re-chunking is supported.pdp-submitskips an aggregate:sub-piece(s) below provider min piece size. In the single-asset path each source CID is its own sub-piece, and the provider enforces a minimum piece size (commonly 1 MiB). CIDs whose CAR pads below that floor cannot go through the passthrough path on that provider — use the multi-asset path (plan --no-auto-packthenpack-cars --pack-target-sizeat or above the provider minimum) to batch them into a large enough piece.- Submission pauses on
spike. The network base fee is above--max-base-fee.pdp-submitwaits out the congestion; check withipfs2foc gas.
docs/personas.md covers per-profile failure modes (gateway
flakes, disk pressure, idle-timeout cascades) and recovery.
- commP pass. For each CID, fetch its CAR (
?format=car&dag-scope=all) from a trustless gateway and stream it through the Filecoin piece hasher to get its PieceCID v2 (FRC-0069). The CAR is rooted at the original CID, so storing it keeps the CID intact, and the CAR root is checked against the requested CID. - Pack. Group source CIDs into sub-pieces and bin-pack those into aggregates by
the
--piece-sizetarget (default 32 GiB; cap it to the provider's maximum piece size). In the single-asset path (plan's default), each source CID becomes one passthrough sub-piece whose pull source is the gateway URL directly; no CAR file touches migrator disk. In the multi-asset path (plan --no-auto-packfollowed bypack-cars), source CIDs are concatenated into assembled sub-pieces — one multi-root CAR file per sub-piece under--car-store. Either way, each aggregate's root is the aggregate piece commitment — the merkle root of its sub-piece commitments, ordered largest-padded-first and zero-padded to the next power of two, the same value the provider re-derives on add. - Pull. For each sub-piece, ask the provider via PDP pull to
POST /pdp/piece/pullfrom<source-base>/piece/{pcidv2}.redirect-servelooks the sub-piece up locally and serves a passthrough sub-piece as a 302 to the gateway CAR or an assembled sub-piece as a byte-served local CAR file. The provider downloads it, verifies its CommP against the declared PieceCID, and parks it. - Aggregate-add.
POST /pdp/data-sets/{id}/pieceswith the parked sub-pieces. The provider recomputes the aggregate piece commitment, confirms it equals the submitted root, and lands one on-chain AddPieces. With the data set'swithIPFSIndexingset, the provider indexes each parked CAR's blocks, so every original CID stays retrievable from the IPFS network by the same CID.
The on-chain operation count is about total_size / piece_size rather than one per
CID, and no payload bytes pass through your machine beyond the single commP read.
A provider's PDP pull admits only source URLs shaped /piece/{pieceCidV2}, and it
follows cross-origin redirects (re-validating scheme and public host). A
/piece/{pcidv2} endpoint that 302s to /ipfs/{cid}?format=car lets the provider
pull the CAR straight from the gateway. The provider verifies pulled bytes against
the PieceCID you supply, so the commP pass (step 1) runs regardless.
The aggregate root is the aggregate piece commitment: the trunc-254 merkle of the
sub-piece commitments, largest-first, zero-padded to the next power of two. The same
value is recomputed by Curio (commputils.PieceAggregateCommP, go-commp-utils) on
add. packages/core/src/piece-aggregate.ts computes it locally so the on-chain add validates; the
add rejects a mismatched root, so a successful commit confirms the local computation.
This value is verified byte-for-byte against go-commp-utils in test/.
Each aggregate moves through planned → submitted → parked → committed (or
failed). parked means the provider has downloaded and verified every sub-piece but
nothing is on-chain yet. pdp-submit caps the count of aggregates at submitted/parked
that have not reached committed (--max-in-flight), so a provider is not asked to
download far more than is then committed, and it pauses when the network base fee is above
--max-base-fee.
Repacking touches only planned aggregates. Once an aggregate is submitted or beyond,
its index and members are frozen, and its CIDs are excluded from future packing.
redirect-serve needs a public HTTPS URL resolving to a public IP. Two
built-in paths:
--ingress cloudflared— spawns a Cloudflare quick tunnel (*.trycloudflare.com). No account, works behind CGNAT, requires thecloudflaredbinary on PATH.--ingress funnel(default) — you run the local HTTP server and front it yourself with Tailscale Funnel, Cloudflare Tunnel, or a VPS reverse proxy.
Setup details, prerequisites, and the public-HTTPS shape the provider
validates live in docs/ingress.md. Pass the HTTPS
origin only (scheme + host, no path) as --source-base.
The serve daemon answers the same /piece/{pcidv2} route, so one process
can carry the console and the pull source: ipfs2foc serve --ingress cloudflared spawns the tunnel and self-checks reachability, or front the
serve port yourself and pass --public-base https://<host>. The console
shows the public piece endpoint and whether it answers.
Two wallets spend on a migration, in different currencies.
The storage provider submits and pays the FIL gas for the on-chain transactions in
this flow: data set creation, AddPieces, and the recurring proof-of-possession
transactions. The migrator authorizes each by an EIP-712 signature carried in the call's
extraData, and the provider sends the transaction.
The migrator is the data set's FWSS payer and spends both currencies:
- USDFC for storage. Data set creation opens a payment rail from the migrator to the
provider and requires the migrator to have deposited enough USDFC to cover the minimum
lockup plus a one-time sybil fee; AddPieces raises the rail's locked amount as the data
set grows. See
FilecoinWarmStorageService.dataSetCreated/piecesAddedin filecoin-services. - FIL for the migrator's own setup transactions, sent from the migrator's wallet: approving USDFC to the FilecoinPay contract, depositing USDFC, and approving FilecoinWarmStorageService as a payments operator with sufficient rate and lockup allowance.
Filecoin gas cost scales with the block base fee, and PDP transactions burn a large
amount of gas, so network congestion multiplies the provider's cost. The gas command
reads the latest block base fee (attoFIL/gas; floor 100) and reports a level: ok,
rising, or spike. Above --max-base-fee (default 1,000,000) the level is spike, and
pdp-submit pauses so submission waits out congestion.
State lives in the SQLite database (migrate.db by default): each CID's piece commitment
and status, the sub-piece (passthrough or assembled) it belongs to, the aggregate plan,
and per-aggregate lifecycle (data set id, transaction hash). A run resumes from here;
re-running plan computes only CIDs that are not yet done, retries failures, and
appends new sub-pieces and aggregates without disturbing prior planning state.
Tables: pieces, sub_pieces, sub_piece_members, aggregates,
aggregate_members.
- The source must serve deterministic trustless CARs. Use
probeto check. - Sub-piece size: each CID's CAR must be within the provider's pull piece limit
(
PieceSizeLimit, ~1 GiB raw).plandoes not check this limit; a CAR larger than the pull cap completesplanand then fails atpdp-submitpull time. Hold large CIDs out of the run until the migrator supports re-chunking. - Aggregate piece size:
--piece-sizeis the target aggregate piece size, bounded by the provider's maximum piece size (up to 64 GiB). A piece whose padded size exceeds the configured--piece-sizeis reported asoversizedand not packed, never silently dropped; this is the aggregate-budget bound, not the pull cap. - Sub-pieces per pull request: the pull admission
eth_call-simulates AddPieces over the batch, and the PDPVerifierPiecesAddedevent carries one piece CID per piece. The FVM caps a single actor event at 8192 bytes, so a pull batch of too many sub-pieces reverts admission.pdp-submitsplits the pull into batches (--pull-batch, default 32), each with its own authorization. The on-chain aggregate-add stays one top-level piece, so the cap applies to the pull batch, not to how many sub-pieces an aggregate holds. - Determinism: the provider re-fetches the gateway CAR and recomputes CommP, so a byte
difference is a permanent failure.
planpins one gateway origin and request shape per piece, andredirect-serve302s to that exact URL. - All-or-nothing aggregate: one unretrievable sub-piece fails its aggregate. The commP pass validates per-CID retrievability first.
The docs/ folder is organized by Diátaxis:
- Tutorial — your first migration on calibration, one CID end-to-end with a checkpoint at every step.
- How-to — the hosted console, the local console, operator profiles (disk/bandwidth/time budgets → knob settings, failure modes, recovery), choosing a gateway, public ingress.
- Reference — command reference and glossary.
- Explanation — how a migration lands on chain (the invariants for integrators), how it works, gas and payments, scope and limits.
- Concurrent pull batches.
pdp-submitpulls sub-piece batches one after another. The provider's pull is the throughput floor, so overlapping batches up to--max-in-flightshortens a large run. - Per-run performance report. The CLI logs stage throughput (commP MiB/s, provider
pull MiB/s, add confirmation time). Persist these per run and surface them in the
dashboard so an operator can tune
--concurrency,--max-in-flight, and--pull-batchagainst the observed provider pull rate, which dominates a large migration. - Sources without trustless CARs. Retrieve through Helia (bitswap + trustless-gateway block brokers), assemble canonical CARs locally, and host each aggregate over public HTTPS for the provider to pull.
Issues and pull requests welcome at
github.com/SgtPooki/ipfs2foc. See
CONTRIBUTING.md for the dev loop and conventions, and
SECURITY.md for key-handling and on-chain-spend guidance.