ML Data Linguist - AWS AI Data

Amazon Development Center U.S., Inc.
US, MA
$32.7K a year
Full-time
We are sorry. The job offer you are looking for is no longer available.

Amazon Web Services (AWS) is looking for a data associate to help with annotations and data analysis. As part of the AiData Team at AWS you will responsible for delivering high-quality training data to ensure the best performance of the AWS machine learning systems.

Our goal is to produce the highest quality training data in the industry and to delight our customers by improving human language understanding and natural language processing.

AWS Utility Computing (UC) provides product innovations from foundational services such as Amazon’s Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS’s services and features apart in the industry.

As a member of the UC organization, you’ll support the development and management of Compute, Database, Storage, Internet of Things (Iot), Platform, and Productivity Apps services in AWS, including support for customers who require specialized security solutions for their cloud services.

Key job responsibilities

  • Build a thorough understanding of data collection and annotation guidelines and various annotation tools.
  • Annotate text data, identifying linguistic categories based on detailed annotation and adhering to guidelines.
  • Perform annotation related tasks; you participate in data generation, collection and quality assurance tasks
  • Collaborate with other ML Data Linguists to resolve data ambiguities and annotation disagreements.
  • Dive deep into the data to perform qualitative error trend analysis.
  • Provide feedback to Language Engineers and Scientists on tool improvements and annotation processes.
  • Diving deep into issues and implement solutions independently
  • Contribute to process improvements to reduce handling time and improve resource output.
  • Develop a variety of language artifacts crucial for model development such as datasets for training and evaluation.

About the team

The Bedrock team is a team of data linguists who primarily support the training of different models in the AWS generative AI platform.

We are specialized in text-based data annotation, writing for ML model training, and toxic content evaluation. Some of the aspects of ML development that the Bedrock team works with include Responsible AI, Reinforcement Learning from Human Feedback, Supervised Fine Tuning, and Human Content Evaluation.

Our team represents a great array of experience in the field of linguistics, including sociolinguistics, computational linguistics, conversation analysis, syntax-semantics, linguistic typology, ESL and foreign languages, as well as translation.

Diverse Experiences

AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply.

If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.

Why AWS?

Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

Inclusive Team Culture

Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences.

Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness.

Mentorship & Career Growth

We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.

Work / Life Balance

We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture.

When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.

Hybrid Work

We value innovation and recognize this sometimes requires uninterrupted time to focus on a build. We also value in-person collaboration and time spent face-to-face.

Our team affords employees options to work in the office every day or in a flexible, hybrid work model near one of our U.S. Amazon offices.

We are open to hiring candidates to work out of one of the following locations :

Virtual Location - MA

BASIC QUALIFICATIONS

  • Bachelor's degree in a relevant field, such as Linguistics, Communications, a foreign language, or other language or data-related disciplines.
  • 6 months of experience with natural language data labeling, data annotation, linguistic annotation or other forms of data markup, and / or teaching experience.
  • Proficient in Spanish, French, German, Portuguese, Japanese, Korean, or another foreign language.
  • Experience identifying linguistic ambiguity and annotation inaccuracies in data.
  • Ability to strictly adhere to annotation guidelines and identify basic parts of speech.
  • Strong organizational skills and detail-oriented
  • Ability to communicate well and actively listen with other data associates on a team.
  • Ability to deliver high quality results under tight deadlines.
  • Comfortable working in a fast paced, collaborative work environment.
  • Willingness to support several projects at one time, and to accept re-prioritization as necessary.

PREFERRED QUALIFICATIONS

  • 1+ years of experience in the language data annotation.
  • Ability to quickly learn new data annotation guidelines, technical concepts, and softwares.
  • Depth and breadth of knowledge in linguistic theory and / or applied linguistics.
  • Familiarity with common text processing tools.
  • Passion for language, linguistics, human language technology and AI.
  • Familiarity with json, yaml, xml or other forms of text markup.
  • Ability to work in different operating systems (Windows, MacOS, or Linux).
  • Ability to navigate a Unix terminal and use common command line tools
  • Knowledge of Python, Java or any other scripting language is a plus.
  • 13 days ago
Related jobs
Promoted
Novartis Group Companies
Cambridge, Massachusetts

The AI & Computational Science (AICS) team at Novartis Biomedical Research is dedicated to revolutionizing drug discovery through AI. Director of Data Science & AI Research Chemistry. AI & Computational Sciences, Novartis Biomedical Research Leadership Role in Chemistry-based AI/Machine Learning. AI...

Promoted
RWS
Wakefield, Massachusetts

General AI Data Annotator - Thai. RWS Group is looking for US-based General Data Annotators to review and score existing AI-generated responses to prompts (questions) in Thai and provide constructive feedback on the responses. General Data Annotator for AI Models - US Only - Thai - Remote, Part Time...

Promoted
My3Tech
Boston, Massachusetts
Remote

Job Title: AWS Data Lake Technical Lead. Client currently uses a Computer Aided Dispatch Records Management System (CAD RMS) system that is currently running in a MySQL instance that will be used in the first phase of the Data Lake build. Design and implement the Clients’ Data Lake in the AWS Enviro...

Promoted
RWS
Wakefield, Massachusetts
Remote

If you already registered with our RWS TrainAI Community and you meet all the requirements, we will reach out to you via email with further details. Data Annotator for AI Models | English (US) | Remote, Part Time, Work from Home. RWS Group is looking for Data Annotators to annotate, label, or tag ...

Promoted
Novartis Group Companies
Cambridge, Massachusetts

Act as an Imaging and data science subject matter expert, to help commodify AI and data analytics expertise and assets into impactful analysis workflows and software products, and to help make our research data FAIR. AI & Computational Science (AICS) is seeking a highly motivated and experienced inc...

Promoted
Diverse Lynx
Boston, Massachusetts

Solid understanding of Databricks fundamentals/architecture and have hands on experience in setting up Databricks cluster, working in Databricks modules (Data Engineering, Client and SQL warehouse). Develop Data Engineering and Client pipelines in Databricks and different AWS services, including S3,...

TetraScience
Boston, Massachusetts
Remote

TetraScience is catalyzing the Scientific AI revolution by designing and industrializing AI-native scientific data sets, which it brings to life in a growing suite of next generation lab data management products, scientific use cases, and AI-based outcomes. TetraScience combines the world's only ope...

Accenture
Boston, Massachusetts

Understand challenges in cloud data migration, building analytical data warehouses and AI/ML use cases on the data platform. Minimum of 8 years of experience in selling Cloud based data solutions, analytical data warehouses, cloud data migration solutions, analytics/reporting. To accelerate our cust...

Data Annotation
Boston, Massachusetts

Join our team to help train AI chatbots while gaining the flexibility of remote work and choosing your own schedule. DataAnnotation is committed to creating quality AI. We are looking for a professional content writer and copy editor to join our team and teach AI chatbots. Projects are paid hourly, ...

Dana-Farber Cancer Institute
MA, US
Remote

The Department of Informatics and Analytics at Dana-Farber Cancer Institute seeks a motivated and talented Artificial Intelligence and Data Engineering Intern for our expanding AI & data team. We are currently seeking to fill 4 internship position: one in computer vision, one in AI data engineer...