Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I really admire what the Internet Archive does, and I'd be very willing to donate to them, but for one caveat: IA's robots.txt policy retroactively applies the current robots.txt on a particular domain to the entire archive of previous captures for that domain. This means that if a website goes offline, and an unrelated third party later acquires the domain, and uses a new, restrictive robots.txt, then the older site is no longer accessible in the archive.

I know this may seem somewhat trivial, but I've run into this problem more than once, and I find that it undermines the value of the archive: if you're trying to preserve history, making that preservation contingent on the the state of things in the present defeats your purpose.

I'd much rather see them adopt a more sensible policy of obeying whatever robots.txt is contemporaneous to each particular site capture.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: