Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

That's not really true - John Resig hooked Rhino up to a good-enough DOM model in a weekend:

http://ejohn.org/blog/bringing-the-browser-to-the-server/

No, you're not going to match Google's crawling infrastructure or data extraction libraries. But if you just want to grab pages off the web, parse them, and handle JavaScript from those pages correctly, you can easily rig something up between Mechanize/html5lib/V8 or Nutch/Tika/Rhino.



Crawling a couple of pages is different from crawling the entire web on a recurring basis. It was hard enough without the emergence of Javascript-enabled pages.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: