python -m scraper.build is the single entry point. It produces, from
scratch, every time:
web/data/jobs.json, web/data/companies.json, web/data/taxonomy.jsonweb/job/<slug>.html — one SEO page per job (JobPosting + FAQPage + Breadcrumb JSON-LD)web/company/<slug>.html — one page per companyweb/category/<slug>.html — one page per career-path/track (AI/ML, DevOps,
Government Jobs, Remote, Freshers, Internships, etc.) — these exist even
with 0 matching jobs yet, with “coming soon” messaging, so nav links
and sitemap entries are never dead on day oneweb/sitemap.xml, web/robots.txtThis was verified end-to-end with a full mocked pipeline run (fake companies across every category/track, checked every generated file exists and every field is correct) before shipping.
.github/workflows/daily_job_post.yml now:
main — so your very first push already
triggers a full build.python -m scraper.build, commits the generated files back
(tagged [skip ci] so it doesn’t loop), and deploys web/ straight
to GitHub Pages in the same run — no separate deploy step needed.GitHub requires a one-time manual toggle, done once, ever: Settings → Pages → Source → “GitHub Actions.” This is a GitHub account/repo setting, not something a workflow file can flip on your behalf the first time. After that one click, every future push/cron/manual run deploys automatically with zero further steps.
Optional: add GROQ_API_KEY as a repo secret (Settings → Secrets and
variables → Actions) if you want AI summaries/skills/FAQs/salary
estimates/eligibility notes. The pipeline runs completely fine without it
— those fields are just left for the cheaper heuristics to fill instead.
scraper/
config.py <- company tokens (Greenhouse/Lever/Ashby) — pre-populated
with ~25 real companies chosen for SEO/traffic reach;
unverified against the live APIs (this build environment
can't reach them), so check the Action's logs after the
first run and prune any that 404.
connectors/ <- one file per source, isolated so one dead source never
kills the run
dedupe.py <- exact + fuzzy duplicate folding
linkcheck.py <- verifies every apply_url, archives dead links after
repeated failures
quality.py <- deterministic 0-100 trust/quality score
eligibility.py <- heuristic "can an India-based applicant apply?" label,
refined by AI when available — never a guarantee,
always shown with a "verify with employer" caveat
taxonomy.py <- deterministic category + cross-cutting track classifier
(AI/ML, DevOps, Government Jobs, Remote, Freshers, ...)
ai_enrich.py <- ONE batched Groq call per NEW job, cached forever
render.py <- pre-renders every job/company/category page + sitemap
build.py <- orchestrates all of the above — the single entry point
web/
index.html <- homepage: hero, category grid, search, filters
(including India-Eligible), dark mode, bookmarks,
skeleton loading — reads web/data/jobs.json client-side
job/, company/, category/ <- generated, individually indexable pages
assets/style.css <- shared theme (light + dark)
data/ <- generated JSON (jobs/companies/taxonomy)
sitemap.xml, robots.txt, .nojekyll <- generated / static
archive/ <- old blog-digest pipeline, kept for reference only
wa.me/918851233153). The Google Form link wasn’t provided, so it’s
intentionally left out rather than shipping a dead link — add a second
<a> next to it in web/index.html’s nav once you have the form URL
(marked with a comment at that exact spot).Account-based bookmarks/alerts across devices, resume/ATS match scoring, AI career-assistant chat, admin moderation UI, analytics dashboard.