studyjob-bot
Our crawler
studyjob.ai indexes public information so that students can see which student jobs fit their degree. This page says exactly what our crawler does, so site operators can decide how to treat it.
- User agent
- studyjob-bot/1.0 (+https://studyjob.ai/bot; ops@studyjob.ai)
- What it fetches
- University programme structure pages and curriculum PDFs (course titles, codes, ECTS, semesters, prerequisites, programme objectives). Course descriptions are read once per semester as extraction input and are never republished. Job postings are collected from public listings and always link back to the original.
- How politely
- At most one request every five seconds per host, a single connection, and standard back-off on 429 or 503. robots.txt is re-read regularly and honoured per path. The crawler never logs in, never uses personal accounts and never circumvents access controls or CAPTCHAs.
- Personal data
- Staff names on course pages are dropped on ingest. Recruiter and job-poster metadata is stripped from postings before storage. Individuals are never shown; only postings, programmes and aggregates.
- Opting out
- Add a Disallow rule for studyjob-bot in your robots.txt, or email ops@studyjob.ai. Opt-out requests are applied within 24 hours and the affected content is removed.