Web Crawler-Based Archiving System
In collaboration with DenktMit, I developed a web crawler-based archiving system to support the intranet migration of a Swiss software manufacturer:
- Developed a web crawler in Kotlin and Selenium for extracting pages with rendered JavaScript
- Implemented a React web app as a search and entry portal
- Set up and optimized a search index using Apache Solr
Technologies
- Kotlin
- Apache Solr
- Selenium
- React
- Docker
- Maven
- Git