Talent.com
Senior Data Acquisition Engineer

Senior Data Acquisition Engineer

People Data LabsSan Francisco, CA, US
30+ days ago
Job type
  • Permanent
Job description

Job Description

Job Description

Note for all engineering roles : with the rise of fake applicants and AI-enabled candidate fraud, we have built in additional measures throughout the process to identify such candidates and remove them.

About Us

People Data Labs (PDL) is the provider of people and company data. We do the heavy lifting of data collection and standardization so our customers can focus on building and scaling innovative, compliant data solutions. Our sole focus is on building the best data available by integrating thousands of compliantly sourced datasets into a single, developer-friendly source of truth. Leading companies across the world use PDL's workforce data to enrich recruiting platforms, power AI models, create custom audiences, and more.

We are looking for individuals who can balance extreme ownership with a "one-team, one-dream" mindset. Our customers are trying to solve complex problems, and we only help them achieve their goals as a team. Our Data Engineering & Acquisition Team ensures our customers have standardized and high quality data to build upon.

You will be crucial in accelerating our efforts to build standalone data products that enable data teams and independent developers to create innovative solutions at massive scale. In this role, you will be working with a team to continuously improve our existing datasets as well as pursuing new ones. If you are looking to be part of a team discovering the next frontier of data-as-a-service (DaaS) with a high level of autonomy and opportunity for direct contributions, this might be the role for you. We like our engineers to be thoughtful, quirky, and willing to fearlessly try new things. Failure is embraced at PDL as long as we continue to learn and grow from it.

What You Get to Do

  • Use and develop web crawling technologies to capture and catalog data on the internet
  • Support and improve our web crawling infrastructure
  • Structure, define, and model captured data, providing semantic data definition and automate data quality monitoring for data that we crawl
  • Develop new techniques to increase speed, efficiency, scalability, and reliability of web crawls
  • Use big data processing platform to build data pipelines, publish data, and ensure the reliable availability of data that we crawl
  • Work with our data product and engineering team to design and implement new data products with captured data, and enhance and improve upon existing products

The Technical Chops You'll Need

  • 7+ years industry experience with clear examples of strategic technical problem solving and implementation
  • Strong software development architecture and fundamentals for backend applications
  • Solid understanding of browser rendering pipeline, web application architecture (auth, cookies, http request / response)
  • Solid programming experience : strong grasp of object-oriented design and experience building applications using asynchronous programming paradigms (e.g., async / await, event loops, or concurrency libraries)
  • Experience building crawlers
  • Proficient in Linux / Unix command line utilities, Linux system administration, architecture, and resource management
  • Experience evaluating data quality and maintaining consistently high data standards across new feature releases (e.g., consistency, accuracy, validity, completeness)
  • People Thrive Here Who Can

  • Must thrive in a fast paced environment and be able to work independently
  • Can work effectively remotely (able to be proactive about managing blockers, proactive on reaching out and asking questions, and participating in team activities)
  • Strong written communication skills on Slack / Chat and in documents
  • You are experienced in writing data design docs (pipeline design, dataflow, schema design)
  • You can scope and breakdown projects, communicate and collaborate progress and blockers effectively with your manager, team, and stakeholders
  • Some Nice To Haves

  • Degree in a quantitative discipline such as computer science, mathematics, statistics, or engineering
  • Experience as a Red Teamer
  • Experience working in data acquisition
  • Experience in network architecture and how to debug and inspect network traffic (DNS, IPv4, Proxies, Application ports and interfaces; packet capture and analysis)
  • Experience with Apache Spark
  • Experience with SQL, including writing advanced queries (e.g., window functions, CTEs)
  • Experience with streaming data platforms (e.g. Kafka or other pub / sub; Spark streaming or other stream processing)
  • Experience with cloud computing services (AWS (preferred), GCP, Azure or similar)
  • Experience working in Databricks (including delta live tables, data lakehouse patterns, etc.)
  • Knowledge of modern data design and storage patterns (e.g., incremental updating, partitioning and segmentation, rebuilds and backfills)
  • Experience with data warehousing (e.g., Databricks, Snowflake, Redshift, BigQuery, or similar)
  • Understanding of modern data storage formats and tools (e.g., parquet, ORC, Avro, Delta Lake)
  • Our Benefits

  • Stock
  • Competitive Salaries
  • Unlimited paid time off
  • Medical, dental, & vision insurance
  • Health, fitness, and office stipends
  • The permanent ability to work wherever and however you want
  • Comp : $160K - $200K

    People Data Labs does not discriminate on the basis of race, sex, color, religion, age, national origin, marital status, disability, veteran status, genetic information, sexual orientation, gender identity or any other reason prohibited by law in provision of employment opportunities and benefits.

    Qualified Applicants with arrest or conviction records will be considered for Employment in accordance with the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act.

    Personal Privacy Policy for California Residents

    https : / / www.peopledatalabs.com / pdf / privacy -policy-and-notice.pdf

    Create a job alert for this search

    Senior Data Engineer • San Francisco, CA, US

    Related jobs
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    Scale AI, Inc.San Francisco, CA, United States
    Full-time
    Software is eating the world, but AI is eating software.We live in unprecedented times - AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, co...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Software Engineer, Big Data

    Senior Software Engineer, Big Data

    ZipRecruiterPalo Alto, CA, US
    Full-time
    We offer a hybrid work environment.Most US-based positions can also.To actively connect people to their next great opportunity. ZipRecruiter is a leading online employment marketplace.Powered by AI-...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Forward Deployed Engineer

    Senior Forward Deployed Engineer

    VirtualVocationsConcord, California, United States
    Full-time
    A company is looking for a Senior Forward Deployed Engineer, AI (Remote).Key Responsibilities Lead the design, development, and deployment of AI / ML-powered solutions tailored to customer needs A...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Management Engineer

    Senior Data Management Engineer

    VirtualVocationsFremont, California, United States
    Full-time
    A company is looking for a Senior Engineer, Data Management (Remote).Key Responsibilities Build and automate data ingestion, transformation, and aggregation pipelines Conduct complex data analys...Show moreLast updated: 1 day ago
    • Promoted
    Senior Data Engineer (Contract)

    Senior Data Engineer (Contract)

    IntelliPro Group Inc.San Mateo, CA, US
    Full-time
    Senior Data Engineer (Gaming / Social Products).Contract (3 months with strong potential to extend).San Mateo, CA (Hybrid or Onsite Preferred). This role will partner closely with the Data Science &...Show moreLast updated: 30+ days ago
    • Promoted
    Lead Data Engineer

    Lead Data Engineer

    VirtualVocationsConcord, California, United States
    Full-time
    A company is looking for a Lead Data Engineer to design, build, and manage enterprise-grade data pipelines.Key Responsibilities Design, develop, and optimize metadata-driven data pipelines in Fab...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    People Data LabsSan Francisco, CA, US
    Permanent
    Note for all engineering roles : with the rise of fake applicants and AI-enabled candidate fraud, we have built in additional measures throughout the process to identify such candidates and remove t...Show moreLast updated: 30+ days ago
    • Promoted
    • New!
    Senior Data Engineer II

    Senior Data Engineer II

    VirtualVocationsHayward, California, United States
    Full-time
    A company is looking for a Senior Data Engineer II to join their data engineering team.Key Responsibilities Design, develop, and maintain scalable data pipelines using Apache Spark on Databricks ...Show moreLast updated: 16 hours ago
    • Promoted
    Senior Cloud Engineer

    Senior Cloud Engineer

    VirtualVocationsConcord, California, United States
    Full-time
    A company is looking for a Senior Cloud Engineer.Key Responsibilities Architect, deploy, and manage EKS clusters on AWS, ensuring scalability, reliability, and cost efficiency Design and impleme...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    VisaFoster City, CA, United States
    Full-time
    Visa is a world leader in payments and technology, with over 259 billion payments transactions flowing safely between consumers, merchants, financial institutions, and government entities in more t...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Software Engineer

    Senior Data Software Engineer

    PsiQuantumPalo Alto, CA, United States
    Full-time
    Quantum computing holds the promise of humanity's mastery over the natural world, but only if we can build a.PsiQuantum is on a mission to build the first real, useful quantum computers, capable of...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    Plum IncSan Francisco, CA, US
    Full-time
    PLUM is a fintech company empowering financial institutions to grow their business through a cutting-edge suite of AI-driven software, purpose-built for lenders and their partners across the financ...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    Prima MenteSan Francisco, CA, US
    Full-time
    Prima Mente’s goal is to deeply understand the brain, to protect the brain from neurological disease and enhance the brain in health. We do this by generating our own data, building brain foun...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    HCLTechSan Jose, CA, US
    Full-time
    We are looking for a highly talented and self- motivated Data Scientist to join us on our journey in advancing the technological world through innovation and creativity. Job Title : Data Engg (Strong...Show moreLast updated: 1 day ago
    • Promoted
    Senior SAP BW Data Engineer

    Senior SAP BW Data Engineer

    VirtualVocationsConcord, California, United States
    Full-time
    Key Responsibilities Implement updates to SAP BW extractors, transformations, info packages, and process chains Develop modern data solutions using Azure Synapse Workflows, Data Pipelines, and S...Show moreLast updated: 2 days ago
    • Promoted
    Senior Data Architect

    Senior Data Architect

    VirtualVocationsFremont, California, United States
    Full-time
    Key Responsibilities : Lead the definition and execution of the Information Management roadmap Manage architectural runway for enterprise data models and advanced data analytics capabilities Pro...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Integration Engineer

    Senior Data Integration Engineer

    VirtualVocationsHayward, California, United States
    Full-time
    Key Responsibilities Integrate and maintain interoperability solutions for clinical and scheduling data from EHR systems Provide operational support during new client integrations and document b...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    VirtualVocationsFremont, California, United States
    Full-time
    A company is looking for a Senior Data Engineer to lead the data engineering function and enhance data capabilities.Key Responsibilities Design, build, and optimize data architecture, pipelines, ...Show moreLast updated: 30+ days ago