We started our second broad crawl for 2019, first step with a byte limit of 50 MB per domain. Second step will have a byte limit of 16 GB per domain. We adjusted the byte limits after having analyzed last year’s broad crawls.
“Ultra big sites”, “OAI-extraction (research databases)”, “Ministries and administrative bodies” and YouTube crawls are running simultaneously with the broad crawl.
We started a selective crawl for the approaching parliamentary elections. We hoped that they would take place together with the elections for the European Parliament, but they will not. At the monthly curator meeting tomorrow, we will have to agree on how to deal with the European parliament elections.
We prepared our list of politicians Facebook profiles and fixed the URL’s as BNE does for the crawl with our Archive-IT account.
We upgraded UMBRA in the production system
We used to start our annual broad crawl on April but due to the changes in our IT Team we had to postpone it until summer.
For this reason, in addition, we have not been able to proceed with BCWeb updating not even study the implementation of UMBRA
We are working intensively in the collections about the diferents elections: European Parlament elections, local elections and Spanish Government elections. We launch a daily harvest in each collection to collect as much as possible of Twitter and Facebook profiles.