jobs_ankurai

AnkurAI Jobs — Static Job Portal

One command builds everything

python -m scraper.build is the single entry point. It produces, from scratch, every time:

This was verified end-to-end with a full mocked pipeline run (fake companies across every category/track, checked every generated file exists and every field is correct) before shipping.

What happens automatically after you push this to GitHub

.github/workflows/daily_job_post.yml now:

  1. Runs on every push to main — so your very first push already triggers a full build.
  2. Also runs daily at 9:00 AM IST, and on-demand from the Actions tab.
  3. Runs python -m scraper.build, commits the generated files back (tagged [skip ci] so it doesn’t loop), and deploys web/ straight to GitHub Pages in the same run — no separate deploy step needed.

The one thing that genuinely can’t be automated from inside the repo

GitHub requires a one-time manual toggle, done once, ever: Settings → Pages → Source → “GitHub Actions.” This is a GitHub account/repo setting, not something a workflow file can flip on your behalf the first time. After that one click, every future push/cron/manual run deploys automatically with zero further steps.

Optional: add GROQ_API_KEY as a repo secret (Settings → Secrets and variables → Actions) if you want AI summaries/skills/FAQs/salary estimates/eligibility notes. The pipeline runs completely fine without it — those fields are just left for the cheaper heuristics to fill instead.

Architecture

scraper/
  config.py       <- company tokens (Greenhouse/Lever/Ashby) — pre-populated
                     with ~25 real companies chosen for SEO/traffic reach;
                     unverified against the live APIs (this build environment
                     can't reach them), so check the Action's logs after the
                     first run and prune any that 404.
  connectors/      <- one file per source, isolated so one dead source never
                      kills the run
  dedupe.py        <- exact + fuzzy duplicate folding
  linkcheck.py     <- verifies every apply_url, archives dead links after
                      repeated failures
  quality.py       <- deterministic 0-100 trust/quality score
  eligibility.py   <- heuristic "can an India-based applicant apply?" label,
                      refined by AI when available — never a guarantee,
                      always shown with a "verify with employer" caveat
  taxonomy.py      <- deterministic category + cross-cutting track classifier
                      (AI/ML, DevOps, Government Jobs, Remote, Freshers, ...)
  ai_enrich.py     <- ONE batched Groq call per NEW job, cached forever
  render.py        <- pre-renders every job/company/category page + sitemap
  build.py         <- orchestrates all of the above — the single entry point

web/
  index.html        <- homepage: hero, category grid, search, filters
                        (including India-Eligible), dark mode, bookmarks,
                        skeleton loading — reads web/data/jobs.json client-side
  job/, company/, category/  <- generated, individually indexable pages
  assets/style.css   <- shared theme (light + dark)
  data/               <- generated JSON (jobs/companies/taxonomy)
  sitemap.xml, robots.txt, .nojekyll  <- generated / static

archive/             <- old blog-digest pipeline, kept for reference only

“Attract as many visitors as possible” — what’s already built toward that

Deferred (still needs a real backend, per your earlier decision to stay static)

Account-based bookmarks/alerts across devices, resume/ATS match scoring, AI career-assistant chat, admin moderation UI, analytics dashboard.