Page tree

Versions Compared


  • This line was added.
  • This line was removed.
  • Formatting was changed.



We are still working on the reorganization of the selective crawls: the strategy is

  • Extension of the selective crawls and smaller broad crawls –
    • We now collect all national Danish news media selectively – both newspaper websites and news media only existing online.
    • We investigate all local news media in order to decide frequency and depth for the future crawls.
    • We made a first crawl of university repositories (with OAI-extraction)

 As Heritrix 3 is not able to archive Facebook profiles. But Archive-IT is able to collect Facebook profiles with an API. We will collect about 100 representative open Facebook profiles at Archive-IT, at the moment we are doing the selection of the profiles.

 We are working on the compression of our archive

We still collect url's for the Olympics event crawl (including the paralympics). We nominate all collected url's for the IIPC collection.



We are continuing to work on this year's broad crawl. We are preparing nas-preload, the tool used to combine the different sources into a single list to be loaded into NAS. This step also includes a DNS check to avoid slowing down the crawl with domains that do not have a DNS response. This year, in addition to excluding domains with no DNS we are also excluding those that give an "unknown" response, as from previous years we know there is generally no content on these domains. Overall the seed list will contain around 4.4 million active domains, and will have improved coverages of the different regional TLDs : .alsace, .paris; .bzh (for Brittany) and the French West Indies.

Turning to project crawls, the 2016 Olympiad is now over but our Olympics crawls are still running. The project, in line with the precedent collaborative collections documenting the 2014 Sotchi Winter Games and 2012 London Summer Games, involves seven curators from the Literature and Art department who work on the selection based on eight themes. Two crawls were planned, before and after the games, covering a list of 558 seeds. Concerning social media, we focused only on Twitter, with 447 French accounts or hashtags collected twice a day from the 4th to the 24th of August. These crawls will be complemented by one for the Paralympic games, to be launched on the 18th of September. We have also communicated our list of seeds for the worldwide collaborative collection led by the British Library for IIPC.


  • we have finally launched our online search interface and would be interested in your feedback. The websites are still not accessible, but it is possible to search for versions either by URL or in our (partial) fulltext. We built a bookmarking feature which allows to save versions online and recall them at the library webarchive terminals.
  • At the moment we have ongoing selective crawls and still an event crawl about presidential elections.