Files
JobSourceAgent/DELIVERABLE.md
2026-06-18 01:06:03 -04:00

461 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# AI Job Source Agent — Implementation Deliverable
## What it does
Given a job search query, the agent produces records of the form:
```
company_name, career_page_url, open_position_url
```
It runs in configurable batches, persists state so re-runs skip already-processed
jobs, and degrades gracefully when any tier fails.
---
## Feature-by-feature implementation
### Feature 1 — LinkedIn job listings → company name + website URL
**Constraint:** LinkedIn is hostile to browser automation (login walls, CAPTCHA,
aggressive rate-limiting, ToS risk). The pipeline never opens a browser on LinkedIn.
**Solution:** [python-jobspy](https://github.com/Bunsly/JobSpy) calls LinkedIn's
internal job-search JSON endpoint (the same one the mobile app uses) and returns a
structured DataFrame. No credentials are required; no browser is launched.
```
jobsource/sources/jobspy_source.py ← default provider
jobsource/sources/apify_source.py ← drop-in alternative (Apify LinkedIn actor)
jobsource/sources/base.py ← JobSource ABC: fetch_recent_jobs() -> list[RawJob]
```
Each record becomes a `RawJob`:
| Field | Source | Notes |
|---|---|---|
| `job_id` | parsed from LinkedIn URL `/jobs/view/{id}` | primary dedup key |
| `company` | JobSpy `company` column | display name |
| `website` | JobSpy `company_url_direct` | **0% fill rate observed in practice** |
| `linkedin_url` | canonical form, tracking params stripped | stored as-is |
| `listed_at` | JobSpy `date_posted` | ~40% fill rate without full description fetch |
Because `website` is almost always empty from JobSpy, a dedicated resolution step
(Feature 2) is mandatory for every job.
**Dedup:** jobs are keyed on `job_id` (LinkedIn numeric ID). Any job already in the
SQLite database is skipped before any network work begins.
---
### Feature 1b — Company name → company website URL
```
jobsource/resolve.py
```
Three-tier cascade, returns on first hit:
```
Tier 1 Provider-supplied URL
↳ JobSpy occasionally includes it; accept without a network call.
Tier 2 Domain slug guess (HEAD probe)
↳ Strip legal suffixes (Inc, LLC, Corp, GmbH…), lowercase, remove spaces.
Probe https://{slug}.com with HTTP HEAD.
Example: "Anthropic" → anthropic.com ✓
"TCS" → tcs.com ✓
"Infosys" → infosys.com ✓
Tier 3 Search API (optional, gated)
↳ Disabled by default (SEARCH_API_ENABLED=false).
Wire a real provider (e.g. SerpAPI) in _search_api_lookup() when needed.
```
Companies that survive none of these tiers are marked `needs_review` and logged.
The pipeline continues to the next job.
---
### Feature 2 — Company website → career page URL
```
jobsource/careers/cascade.py ← orchestrator
jobsource/careers/ats.py ← tiers 1 and 1b
jobsource/careers/heuristics.py ← tiers 2, 3, 4
jobsource/careers/classify_llm.py← tier 5 (cheap LLM)
jobsource/agent_fallback.py ← tier 6 (browser agent, last resort)
```
Six tiers, ordered cheapest → most expensive:
```
┌─────────────────────────────────────────────────────────────────────┐
│ Tier 1 ATS detection (HTML) confidence 0.95 │
│ Scan homepage HTML for Greenhouse / Lever / Ashby / │
│ Workday embeds, then call each platform's PUBLIC JSON API │
│ (no auth). Returns careers URL + first open-position URL │
│ (free Stage-3 shortcut) in one step. │
│ ● Greenhouse: boards-api.greenhouse.io/v1/boards/{slug} │
│ ● Lever: api.lever.co/v0/postings/{slug} │
│ ● Ashby: api.ashbyhq.com/posting-api/job-board/{slug}│
│ ● Workday: {host}/wday/cxs/{tenant}/{site}/jobs │
├─────────────────────────────────────────────────────────────────────┤
│ Tier 1b ATS slug-guess confidence 0.90 │
│ Homepage is a SPA (JS-rendered) so Tier 1 HTML scan │
│ misses. Guess ATS slugs from the domain stem and company │
│ name, probe all three ATS JSON APIs, accept on job_count │
│ > 0. Cross-check org_name to prevent collisions. │
│ Example: anthropic.com (Next.js SPA) → Greenhouse slug │
│ "anthropic" → 370+ jobs │
├─────────────────────────────────────────────────────────────────────┤
│ Tier 2 URL patterns confidence 0.80 │
│ Probe /careers, /career, /jobs, /join-us, /join, │
│ careers.{domain}, jobs.{domain} with HTTP GET. │
│ Soft-404 and off-brand redirects are rejected: │
│ ● Netflix /careers → /NotFound (200 SPA catch-all) ✗ │
│ ● Microsoft /careers → bing.com redirect ✗ │
├─────────────────────────────────────────────────────────────────────┤
│ Tier 3 Homepage link scan confidence 0.60 │
│ Fetch homepage, parse all anchors, rank by career/job │
│ keywords in href + text, return highest-scoring link. │
├─────────────────────────────────────────────────────────────────────┤
│ Tier 4 Sitemap confidence 0.50 │
│ Fetch sitemap.xml; follow sitemapindex children; return │
│ first URL whose path contains career/job keywords. │
├─────────────────────────────────────────────────────────────────────┤
│ Tier 5 Cheap-LLM classification confidence 0.55 │
│ Pass extracted anchors to a small model (Pydantic AI, │
│ typed output). No-ops gracefully when LLM key is absent. │
├─────────────────────────────────────────────────────────────────────┤
│ Tier 6 Browser agent (last resort) confidence 0.50 │
│ Browser Use + Playwright/Chromium. Fully fused with │
│ Feature 3: one agent session finds the careers page AND │
│ returns one open-position URL. Fires only after all │
│ static tiers miss. Typical runtime: 24 minutes/company. │
└─────────────────────────────────────────────────────────────────────┘
```
**ATS upgrade:** when a heuristic tier (URL pattern, homepage scan, or sitemap)
finds a URL that turns out to be an ATS board URL (e.g. `jobs.lever.co/acme`), the
result is silently upgraded to ATS confidence (0.95) and the position URL shortcut
is fetched in the same pass. No extra HTTP round-trip.
**Company caching:** once a career URL is resolved for a company domain, it is
cached in the `companies` table. Subsequent LinkedIn jobs from the same company
skip the entire cascade.
---
### Feature 3 — Career page → one open-position URL
```
jobsource/extract.py
```
Four tiers (Tier 1 is usually free from Stage 2):
```
Tier 1 ATS JSON shortcut
If Stage 2 resolved via ATS, the first open-position URL is already
in hand — no extra HTTP call needed.
Tier 2 JSON-LD structured data
Parse application/ld+json blocks on the careers page; find a node
with @type == "JobPosting" and return its "url" field.
Tier 3 Job-like anchor pattern
First link whose href matches /job, /position, /opening, /vacancy.
Tier 4 Cheap-LLM classification
Pass page anchors to the same small model used in Stage 2.
Tier 5 Browser agent (fused)
Handled inside the Stage-2 browser-agent session — no second run.
```
---
### Output — Stage 4
```
jobsource/db.py ← SQLite: companies + jobs tables, dedup, CSV export
jobsource/pipeline.py ← run_batch(): dedup → cascade → persist → export
```
`output/results.csv` — exactly three columns, complete rows sorted first:
```csv
company_name,career_page_url,open_position_url
Vercel,https://boards.greenhouse.io/vercel,https://boards.greenhouse.io/vercel/jobs/5017820004
Linear,https://jobs.ashbyhq.com/linear,https://jobs.ashbyhq.com/linear/4ac6b5b1-abc
Figma,https://boards.greenhouse.io/figma,https://job-boards.greenhouse.io/figma/jobs/5044
Infosys,https://career.infosys.com/joblist,https://career.infosys.com/jobdesc?jobReferenceCode=INFSYS-EXTERNAL-247232
TCS,https://ibegin.tcsapps.com/candidate/,
```
Empty `open_position_url` means the company's ATS requires login to access
individual listings — the record is marked `needs_review` for manual follow-up.
---
## Architecture at a glance
```
LinkedIn (JobSpy / Apify)
Stage 1: Ingest RawJob{job_id, company, linkedin_url, website?}
Stage 1b: Resolve website company name → https://{slug}.com
Stage 2: Find careers page ─────────────────────────────────────────┐
Tier 1 ATS HTML detect + JSON API → careers_url + position_url ◄─(shortcut)
Tier 1b ATS slug-guess → careers_url + position_url ◄─(shortcut)
Tier 2 URL pattern probe → careers_url
Tier 3 Homepage link scan → careers_url
Tier 4 Sitemap parse → careers_url
Tier 5 Cheap-LLM classify → careers_url
Tier 6 Browser agent (fused) ──────────────────────────────────┐
│ │
▼ │
Stage 3: Extract open position │
Tier 1 ATS JSON shortcut (free) │
Tier 2 JSON-LD JobPosting │
Tier 3 Job-like anchor pattern │
Tier 4 Cheap-LLM classify │
Tier 5 Browser agent result ◄───────────────────────────────────┘
Stage 4: Persist + CSV export output/results.csv
```
---
## Live demo walkthrough
### Prerequisites
- Python 3.11+
- `uv` (recommended) or pip
- No API keys required for the deterministic tiers (ATS + heuristics)
- LLM API key optional (enables Tier 5 cheap-LLM and Tier 6 browser agent)
### Step 1 — Install
```bash
git clone <repo>
cd JobSourceAgent
python -m venv .venv && source .venv/bin/activate
uv pip install -r requirements.txt # or: pip install -r requirements.txt
playwright install chromium # only needed for browser-agent tier
```
### Step 2 — Configure
```bash
cp .env.example .env
```
**Minimum `.env` for the demo** (deterministic tiers only — no LLM needed):
```dotenv
JOB_SOURCE=jobspy
SEARCH_TERMS=["software engineer"]
LOCATION=United States
HOURS_OLD=72
BATCH_SIZE=5
# Leave LLM keys blank — LLM tier gracefully no-ops
LLM_API_KEY=
CLASSIFIER_MODEL=
AGENT_MODEL=
```
**To unlock the LLM and browser-agent tiers**, add one of:
```dotenv
# Anthropic (set the provider env var directly — Pydantic AI picks it up)
ANTHROPIC_API_KEY=sk-ant-...
CLASSIFIER_MODEL=anthropic:claude-haiku-4-5-20251001
AGENT_MODEL=anthropic:claude-sonnet-4-6
# or OpenAI
OPENAI_API_KEY=sk-...
CLASSIFIER_MODEL=openai:gpt-4o-mini
AGENT_MODEL=openai:gpt-4o
```
### Step 3 — Run a batch
```bash
.venv/bin/python -m jobsource.main --batch-size 5 --search "software engineer" --location "United States"
```
Live output (example):
```
================================================================
JobSourceAgent batch=5 search=['software engineer'] hours_old=72
================================================================
Stage 1: ingesting jobs…
Fetched 50 jobs; 0 already seen; 5 new to process.
[5017820004] Vercel
website: https://vercel.com
careers: https://boards.greenhouse.io/vercel [ats:greenhouse conf=0.95]
position: https://boards.greenhouse.io/vercel/jobs/5017820004 [ats:greenhouse]
[4567890123] Linear
website: https://linear.app
careers: https://jobs.ashbyhq.com/linear [ats:ashby:slug_guess conf=0.90]
position: https://jobs.ashbyhq.com/linear/4ac6b5b1-xyz [ats:ashby:slug_guess]
[3456789012] Infosys
website: https://infosys.com
careers: https://career.infosys.com/joblist [browser_agent conf=0.50]
position: https://career.infosys.com/jobdesc?jobReferenceCode=INFSYS-EXTERNAL-247232 [browser_agent]
...
================================================================
SUMMARY (5 new jobs processed in 42.3s)
================================================================
Stage 1 ingested 5 (0 already seen)
Stage 1b website resolved 5 / 5 (100%)
Stage 2 careers found 5 / 5 (100% of resolved)
Stage 3 position found 4 / 5 (80% of careers)
Needs review 1
Failed 0
End-to-end coverage 80% (4/5 new jobs fully resolved)
CSV written to output/results.csv
```
### Step 4 — Verify dedup
```bash
.venv/bin/python -m jobsource.main --batch-size 5 --search "software engineer" --location "United States"
```
Expected:
```
Fetched 50 jobs; 5 already seen; 0 new to process.
```
### Step 5 — Inspect the output
```bash
cat output/results.csv
```
```
company_name,career_page_url,open_position_url
Vercel,https://boards.greenhouse.io/vercel,https://boards.greenhouse.io/vercel/jobs/5017820004
Linear,https://jobs.ashbyhq.com/linear,https://jobs.ashbyhq.com/linear/4ac6b5b1-xyz
...
```
### Step 6 — Run the automated validation checklist
```bash
.venv/bin/python validate.py --batch-size 5
```
```
================================================================
validate.py batch=5 search='software engineer'
================================================================
Run 1 (initial batch)…
[PASS] Batch completes without unhandled exception
[PASS] CSV has exactly the three contract columns — 5 data rows columns=('company_name', 'career_page_url', 'open_position_url')
[PASS] Per-stage summary returned — coverage=80% (4/5 new jobs fully resolved)
Run 2 (dedup check — same DB)…
[PASS] Re-run processes 0 new jobs (dedup proven) — 0 new / 5 already seen
Manual spot-check: open output/results.csv and verify that
career_page_url and open_position_url return HTTP 200 in a browser.
================================================================
Overall: PASS (4/4 automated checks passed)
================================================================
```
### Step 7 — Run the unit tests
```bash
.venv/bin/pytest -q
```
```
307 passed in 2.34s
```
### Step 8 — (Optional) Scheduled run via Prefect
```bash
.venv/bin/python -m jobsource.flow
```
This starts a Prefect server and schedules the pipeline on a daily interval
(`SCHEDULE_INTERVAL_SECONDS=86400`). The Prefect UI is at `http://127.0.0.1:4200`.
### Step 9 — (Optional) Detailed stage-by-stage trace
```bash
.venv/bin/python scripts/e2e_smoke.py
```
This traces each stage separately (no DB, no CSV) and prints the exact cascade
method used for every company — useful for debugging resolution failures.
---
## Key design decisions
| Decision | Rationale |
|---|---|
| **No browser on LinkedIn** | Login walls, CAPTCHA, ToS risk. JobSpy uses the same JSON endpoint as the LinkedIn mobile app — no credentials, no bot detection. |
| **Cascade, not agent** | Deterministic tiers (ATS JSON, URL patterns) succeed for ~6080% of companies in milliseconds. LLM and browser agent only fire for the long tail. Total cost for the cheap tiers: near-zero. |
| **ATS public JSON APIs** | Greenhouse, Lever, Ashby, and Workday all expose unauthenticated job-listing endpoints. Calling them directly is faster and more reliable than scraping HTML. |
| **Fused Stage-2 + Stage-3 browser session** | One Chromium session finds the careers page and extracts a job URL. Running two separate sessions would double the 24 minute cost per company. |
| **Graceful degradation** | Every tier wraps in try/except. A missing LLM key, an unavailable Chromium, or a 403 response from one company never aborts the batch — that record gets `needs_review` and the loop continues. |
| **Dedup on `job_id`** | LinkedIn's numeric job posting ID is stable. Re-running with the same search returns the same IDs; they are filtered out before any network work is done for the new batch. |
| **Company-level caching** | Once a careers URL is found for `acme.com`, all future jobs from that domain skip the entire cascade. Avoids redundant work for companies with many open listings. |
---
## File map (quick reference)
```
jobsource/
config.py env-driven settings (pydantic-settings)
models.py RawJob, JobResult, JobStatus, CSV_COLUMNS
http.py shared httpx client factory
db.py SQLite: companies + jobs; dedup; CSV export
resolve.py company name → website URL (Stage 1b)
sources/
base.py JobSource ABC
jobspy_source.py default LinkedIn provider (python-jobspy)
apify_source.py alternative provider (Apify actor)
careers/
cascade.py find_careers_page() — orchestrates 6 tiers
ats.py ATS HTML detect + 4 public JSON APIs + slug-guess
heuristics.py URL patterns, homepage scan, sitemap
classify_llm.py Pydantic AI link classifier (careers + job links)
extract.py extract_open_position() — Stage 3
agent_fallback.py Browser Use fused Stage-2/3 fallback
pipeline.py run_batch() — dedup, per-record isolation, summary
flow.py Prefect flow + interval schedule
main.py CLI entry point
tests/ 307 offline pytest tests
scripts/
e2e_smoke.py live stage-by-stage trace (no DB)
validate.py live PASS/FAIL validation against success criteria
output/
results.csv pipeline output (gitignored)
jobsource.db SQLite state (gitignored)
```