post: split the six months post into a series of five
All checks were successful
Deploy to OVH VPS / deploy (push) Successful in 52s

This commit is contained in:
Lorenzo Iovino 2026-09-29 00:17:42 +02:00
parent c6b3007995
commit 015f4b0127
5 changed files with 248 additions and 80 deletions

View file

@ -0,0 +1,77 @@
---
title: "Italian public notice boards are a zoo"
description: "About 100 vendor platforms, PDFs that are not PDFs and a municipal domain that now sells Viagra. What I found crawling every albo pretorio in Italy."
pubDate: 2026-10-02
tags: ["scraping", "rust", "open-data", "public-administration", "privacy"]
---
In the [first post of this series](/blog/six-months-reading-italian-municipalities) I told how a small Rust exercise on one town became a crawler for every Italian municipality.
This one is about the part nobody sees: **how** you read 9,617 public notice boards when there is no standard, no API and no feed.
Short answer: one vendor at a time. 😅
## The scraper that became a zoo
The first scraper was for one vendor platform. Easy. Then I added a second town and it was on a different platform. Then a third one.
At some point I stopped counting towns and started counting **vendors**. Today there are about **100** of them in the codebase, and each one has a personality:
- old Java server pages with a hidden "ViewState" that you must send back exactly as you received it, or nothing works
- sites that only work inside a real browser, so the crawler drives a headless Chromium
- PDFs that say they are PDFs but they are signed `.p7m` files
- pages that lie about their encoding
- "search" pages that return everything if you pass `limit=0` (thank you, I'll take it 🙏)
The good news is that vendors are few compared to towns. When you understand one platform, you often get hundreds of municipalities for free, because they all bought the same software. So the real map is not "7,896 towns", it's "who sells software to the public sector".
## Who owns the notice board?
One number surprised me: **67.6%** of Italian municipalities publish their notice board on a domain they don't own. It's the vendor's domain.
<br/>
The official record of what a town decided lives on a `.com` of a private company.
It works, most of the time. But if the vendor changes, closes or forgets to renew something, the history of that town moves or disappears with it.
## The Viagra domain
And then there is my favorite discovery.
Some public bodies were closed years ago (mountain communities, old consortiums) and nobody renewed their domain. Somebody else bought it.
<br/>
One of them, still listed as "official website" in the national registry of public administrations, today is an **online pharmacy selling Viagra without prescription**. In Italian. 🙃
For a crawler this is a real problem, not only a funny story: the domain answers `200 OK`, the page loads, everything looks alive. You must look at *what* the page says, not only *if* it answers.
Also funny: I tried to add the 20 Regions too, before understanding that most Regions **don't have an albo pretorio** at all. Lesson learned: read the law before writing the crawler.
## PDFs, LLMs and regex
The notice board gives you a title and a PDF. The interesting part is inside the PDF: who gets the money, how much, for what.
So the pipeline got longer:
1. download the PDF (4.6 million so far)
2. extract the text, with OCR when it's a scan, and it's a scan more often than you think
3. pull out the fields: amounts, suppliers, dates, offices
For step 3 I started with a local model with Ollama, then moved to cheaper hosted models. Then I learned the boring lesson: you can do **a lot** with regex before calling any model. A spending approval in Italian has a very recognizable shape. The LLM is for what the regex can't read, not for everything.
Cheaper, faster and, surprise, easier to test.
## The serious part: privacy
These acts are full of personal data. Names, addresses, health situations of people who asked the town for help.
The law says they must stay public for 15 days. That doesn't mean "republish forever on a search engine".
<br/>
So nothing goes online before a privacy filter checks it, and whole categories of acts are never shown at all. Crawling everything is fine. Showing everything is not.
## What I learned
- **There is no standard, only vendors.** If you want to read the public sector, you must understand who sells software to it.
- **A 200 is not a proof of life.** Check what the page says.
- **Regex first, LLM after.** It's not sexy, but it pays the bills.
- **Collect everything, publish carefully.** Two different decisions.
Next time: the festivals hidden inside all these acts. 🎉