Data provenance
How PavedIT builds its career graph
Beta notice · Updated August 8, 2026
Jobs
PavedIT prioritizes employer-authorized feeds, then official public ATS endpoints, then approved structured-data collection. Licensed providers are used only for measured coverage gaps. A source must pass terms review, robots policy where applicable, rate limits, canonical-URL validation, attribution, freshness, duplicate, and removal checks before its jobs can be published.
Taxonomy
PavedIT owns its graph schema, mappings, corrections, and derived evidence—not the public classifications it builds upon. O*NET, ESCO, SOC/BLS, and future licensed sources retain their source URI, version, license, and attribution. PavedIT-derived relationships are labeled separately with their method and evaluation status.
Collection boundaries
PavedIT never bypasses logins, CAPTCHAs, paywalls, access tokens, IP blocks, or other technical restrictions. It does not collect candidate profiles. Approved crawlers identify themselves, provide a monitored contact address, stay within approved hosts, and can be disabled source by source.
Current status
The hybrid activation pipeline supports employer JSON feeds, Greenhouse, Lever, Ashby, and bounded schema.org collection. No source is enabled until its database migration, server-only configuration, approval record, attribution review, and acceptance measurements are completed. An empty catalog remains preferable to unapproved or fabricated data.