
Real job — pulled straight from Zocket’s careers page · Verified July 31, 2026 · No reposts.
Job description
Zocket is hiring a Data Engineer — a full-time, based in Chennai, India role. Apply directly on Zocket's careers page below.
Senior Data Engineer
Location: Chennai, India
Department: Technology
Experience: 4-8
Assessment — Data Engineer (Mid-level), Consumer Research
Crawl a brand across many sources, then pipeline & enrich the data
The objective
What you must deliver
1. Crawl the brand across at least 5 distinct sources
- Work out where meaningful public data about the brand lives, and crawl five or more genuinely different sources. Which sources you choose — and why — is a core part of the evaluation. (Five pages of the same site is not five sources.)
- Handle each source's realities: pagination, rate limits, retries/backoff, and things breaking.
- Make the crawl incremental (a re-run collects only what's new) and replayable (persist raw responses).
- In the design doc: list the sources you picked, why each one, what it contributes, and what you'd add with more time.
2. Pipeline the data into a clean, queryable store
- Unify data from all sources into one coherent model — you decide the schema and the store, and justify both.
- Idempotent & incremental; deduplicate, and handle the same item arriving from multiple sources or changing over time.
- Data-quality checks that fail loudly, plus enough run metrics / logging to actually operate it.
- Structure it so it could run on a schedule, and include a short sketch of how you'd orchestrate it. (No need to deploy.)
3. Enrich the data
- Add the structured signals that make this useful for understanding the brand — at minimum sentiment toward the brand and the themes/topics being discussed.
- Whatever method you use must be reliable inside a pipeline (it handles bad, failed, or malformed results) and evaluable — and you must evaluate it: hand-label a sample, report how well it does, and show where it fails.
4. Answer a question about the brand
- Expose at least one analytical answer to a real question about the brand (you choose the question). Include the query/view and sample output.
5. Design doc (~1–2 pages) — the part we read most closely
Completion Gate (what "done" means)
Deliverables & how to submit
- A Git repository (GitHub link easiest).
- README.md — setup, run, and time spent.
- The design doc and your enrichment evaluation.
- Sample data + outputs committed, or a one-command way to regenerate them.
- A short walkthrough video (≤ 5–7 min): run it, show a re-run staying idempotent, and talk through your design and source choices.
- Optional (noticed): a feature branch + a real PR description.
Ground rules
- Crawl politely and legally: public data only, honor robots/ToS, throttle, no logins or paywalls. If a source is off-limits, choose another and say so.
- Keep volumes modest — this is a design test, not a load test.
- Libraries and AI assistants are fine — but you must be able to defend every decision in the walkthrough.
- Timebox honestly. A tighter, well-reasoned slice beats a sprawling unfinished one. Tell us what you'd do next.
Get Data Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Data Engineer at Zocket?
The required skills for Data Engineer at Zocket include: Python, SQL, ETL, Data Warehousing, NLP, Machine Learning, API, Git, Docker, Cloud Computing.
What is the seniority level for Data Engineer at Zocket?
Data Engineer at Zocket is a Mid Level / Senior level position.
How do I apply for Data Engineer at Zocket?
You can view the full description and apply for Data Engineer at Zocket on EchoJobs: https://echojobs.io/job/zocket-senior-data-engineer-vlvnx.