All integrations
Internet Archive logo

Internet Archive

File StorageResearch & AcademiaMisc References

Search the Internet Archive's library of 40 million free books, films, concerts, images, software and datasets, read any item's metadata and files, download a file, and upload your own: create items, add and remove files, edit metadata, follow the background tasks archive.org runs on them, see how often an item is viewed and where its readers came from, and leave a review.

20 actions

Actions

Steps your workflow can run in Internet Archive.

Search itemsSearch archive.org's 40 million items: books, films, concerts, software, images and datasets. This is the search behind the site's own search box, and it takes the same query language: bare words search everything, 'field:value' narrows to one field, quotes match a phrase, and AND, OR and NOT combine clauses. Reaches the first 10,000 results; use 'Scrape items' to walk a whole result set.
Scrape itemsWalk a whole search result set, however large, a page at a time. Each run returns up to 10,000 items and a cursor; feed that cursor back in, with the same query, to get the next page, and stop when no cursor comes back. Turn on 'Count only' to get the size of a result set without any of it.
Check an identifierCheck whether an identifier is free before creating an item with it. An identifier is an item's permanent name and the last part of its archive.org URL, and it cannot be changed casually afterwards. Also applies archive.org's own rules about what a name may be: 5 to 100 characters, letters, digits, periods, underscores and dashes, which the site's availability check does not.
Get an itemEverything archive.org knows about one item in a single request: its descriptive metadata, the files it holds with their sizes, formats, checksums and download URLs, its total size, when it was last changed, and whether it has been hidden. The embedded file list is capped at 100 entries: use 'List item files' to page through a larger one.
Get one metadata fieldRead one part of an item's record instead of all of it: its size, the date it last changed, its descriptive metadata, or one entry from its file list. Much smaller than 'Get an item' when the item holds thousands of files, and the only way to slice a long list with a start and a count.
List item filesPage through the files inside an item. Each file comes back with the exact name archive.org knows it by, its size, format, checksums and a direct download URL. Filter to the originals to skip the thumbnails, OCR text and transcodes archive.org generates from them, or narrow by part of a name.
Download a fileFetch one file out of an item and keep it as a document a later step can read, attach or send on. The bytes come straight from the data node holding the item. Files up to 200 MB; for anything larger, use the download URL 'List item files' returns.
Create an itemCreate a new item on archive.org and put its first file in it. The identifier becomes the item's permanent name and the last part of its URL, so check it is free first. archive.org accepts the file straight away and ingests it in the background, so the item's page can take a few minutes to appear. Everything uploaded here is public. If the identifier is already taken the step stops rather than quietly dropping the metadata below - archive.org only reads it while an item is being made - unless you turn on 'Add to it if it already exists'.
Upload a fileAdd a file to an item that already exists. Uploading a name the item already holds replaces that file; turn on 'Keep the previous version' to have archive.org move the old one into the item's history instead.
Delete a fileRemove one file from an item. The item itself stays, and so do its other files. Turn on 'Also delete derivatives' to take the thumbnails, OCR text and transcodes archive.org generated from this file with it.
Update item metadataChange the metadata on an item you own: its title, description, subjects, licence or any other field, or on one of its files. Setting a field that is not there adds it; setting one that is replaces it; a field already holding the value you asked for is left alone, so running the step twice is not an error. The change is queued as a background task, and the task id comes back so a later step can watch it - or 'changed' comes back false when there was nothing to do.
List tasksThe background jobs archive.org is running on your items. Every upload, metadata change and review becomes one of these, so this is how a workflow follows what it started. 'Active' holds the queued, running, failed and paused ones; 'Finished' holds the ones that have already run.
Get a task logRead what a task actually did. This is where a failed upload, derive or metadata change says why. Logs run to megabytes on large items, so this returns the last lines, which is where a failure states itself, and reports how many there were in total. A task only has a log once it has started running, so one submitted moments ago comes back with 'available' false and its run state rather than an error: wait, then read it again.
Submit a taskAsk archive.org to run a job on one of your items. Deriving regenerates the thumbnails, transcodes and OCR text after its files change; darkening hides an item from the public and undarkening puts it back; renaming gives it a new identifier. The job is queued rather than run, and the task id that comes back is what 'Get a task log' reads.
Get item view countsHow often items have been viewed: all time, over the last 30 days, and over the last 7. Takes up to 50 identifiers at once, so a whole collection's headline numbers are one step.
Get daily view countsA day-by-day view series for items, split between human traffic, robots and requests archive.org could not classify. One row per date, newest last, ready to chart. archive.org holds this data from 1 January 2017 onward and publishes it with a few weeks' lag.
Get view detailsWhere an item's views came from: a breakdown by country and region with coordinates, split between human and robot traffic, plus the pages that linked to the item over the same window. Give a start and end date for a fixed period, or a number of days for a rolling one.
Get your reviewRead the review your account has left on an item, with its title, body, star rating and the dates it was written and last changed. archive.org's reviews API is scoped to the connected account, so this reads your own review rather than the item's whole review list. Not having written one is not an error: it answers 'has_review: false'.
Add or update a reviewLeave your account's review of an item, with an optional star rating. One account has one review per item, so running this again replaces what was there. The review is queued as a background task and appears on the item's page once that task runs.
Delete your reviewRemove your account's review of an item. Deleting a review that was never written is not an error: it answers that there was nothing to remove.

Connect in a few clicks

Authenticate once and every action and trigger for the app is ready to drop into a workflow. No glue code, no maintenance.

Automate across your stack

Chain apps together with triggers, actions, and logic that move data between your tools automatically, so work happens without you.

Secure by default

Credentials are encrypted and scoped per workspace. Connect the tools your team already trusts with confidence.

Automate Internet Archive with Lodol.

Connect Internet Archive and build your first workflow in minutes. No credit card required.

Free plan available · No credit card required