Create an account for powerful AI tools, award-winning courses, and access to our vibrant community.
Already have an account?
Join 250,000+ professionals and teams at Microsoft, Shopify, and even NASA. đ
Already have an account? Login
Find the best remote jobs. Answer a few questions and we'll deploy a powerful assistant to help you search, create alerts, and more.
1 What roles are you open to?
2 Experience level
3 Work style
Did you know? If memory is enabled, Writing.io can remember your job search preferences and help you to improve your resume, craft customized outreach and more.
Category
Evaluate and test AI chatbots through structured conversations and voice interactions to provide feedback that improves model performance and safety.
Productive Playhouse is building a talent pool of Spanish (Spain) speakers for an upcoming project testing and evaluating leading AI chatbots. The objective is to enhance response quality through direct user interaction with the various AI models.
Open to freelancers based outside the U.S.
This is more than one job. Apply once, and weâll keep you in mind for this project â and future ones that need your language skills.
Weâre looking for independent contractor engagement (task-based - project-based).
As an AI Evaluator, youâll play a critical role in shaping and improving the next generation of Generative AI. Youâll participate in structured, hands-on evaluations by interacting directly with various AI models to assess their capabilities, safety, and helpfulness. Your insights and data will directly inform model development and optimization.
Exact tasks vary by project and will be spelled out in that projectâs Statement of Work (SOW) before you start. Depending on the role, your work may include:
đ Your language skills are the whole point â and this project pays for them
⥠Apply once, get contacted directly when a spot opens â no chasing
đ Work from anywhere, on your own schedule, alongside any other work you do
đ¤ Clear task specs, no ongoing oversight, no micromanagement
đź Real project experience with a global data and language services company
⨠Work that actually improves how AI understands your language
We started by teaching kids through award-winning programming. We grew into a global data and language services company trusted by clients worldwide. But the mission never changed: keep language alive.
Transcription, translation, localization, linguistic analysis, AI evaluation â it all comes back to preserving languages and the cultures they carry.
All engagements are contingent upon successful completion of identity verification. Productive Playhouse is an equal opportunity organization headquartered in California and committed to diversity and inclusion across our global workforce. We welcome applicants of all backgrounds and abilities, regardless of location or engagement type. For accommodations or inquiries, please contact [email protected].
Create and evaluate coding tasks for AI agents by building realistic developer environments, designing challenges, and writing tests to assess model performance.
Please submit your CV in English and indicate your level of English proficiency.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves Weâre building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.
Youâll create challenging tasks and evaluation criteria within realistic simulated environments:
What this is NOT
What we look for
Why this is hard
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
How it works
Apply â Pass qualification(s) â Join a project â Complete tasks â Get paid
Compensation
Up to $50/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.
Review AI-flagged student online activity for safety risks, conduct risk assessments, and alert school personnel to potential harm situations in real-time.
As a member of the Securly On-Call team (formerly Securly 24), you are the critical human intelligence that powers our safety mission. Reporting to the Director of Student Safety, you play a vital role in reviewing online activity recognized and âflaggedâ by Securlyâs AI as potentially harmful to students. You will work in tandem with our award-winning technology to perform thorough risk assessments, distinguishing between curiosity and crisis, and executing critical communication protocols to alert school safety teams. This is a high-impact, remote-first role requiring a mission-driven mindset to support student well-being in real-time.
Securly is the market leader in AI-powered student wellness and safety solutions, protecting over 20M students across 20,000+ schools worldwide. Our technology has analyzed over 10B activities to help keep students safe, secure, and ready to learn. We operate with the scale of a tech giant and the agility of a startup, combining innovation, data, and compassion to solve real problems.
Our Impact & Recognition
Our Culture of Engagement
At Securly, culture is our foundation. We are a fully remote, people-first organization built on trust and accountability. Our engagement data far exceeds global benchmarks:
Securly is committed to building a diverse and inclusive workplace. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, disability, or any other legally protected characteristic. Accommodations are available throughout the hiring process. Please contact recruitment.us@securly.com.
#LI-REMOTE #LI-DO1
Provides Korean bilingual audio data and feedback to train and improve AI language models.
Annotates and labels audio data for AI model training, listening to speech samples and marking phonetic details using annotation tools.
Headquarters: 1 Sansome Street, STE 1400 San Francisco, CA 94104 USA
URL: http://www.argosmultilingual.com
We are looking for experienced audio data annotator with US English for our new project!Â
Profile we are looking for:Â
Â
To apply: https://weworkremotely.com/remote-jobs/argos-multilingual-inc-ai-data-annotator-remote
Rates and evaluates advertisement quality to train AI models, providing feedback on ads relevance and compliance.
Uses Avid Media Composer to annotate and label video data for AI model training and development.
Review and rate text, webpages, and images to provide feedback on AI search technology relevance and quality.
Headquarters: Las Vegas, Nevada
URL: https://jobs.telusdigital.com/search/jobs?cfm5=Artificial+Intelligence&ns_category=artificial-intelligence
We are looking for an independent, flexible, remote opportunity where you can help improve AI-powered search technology from the comfort of your home? If you're curious, internet-savvy, and enjoy evaluating online content, this freelance project could be a great fit.
Â
A Day in the Life of a Content Reviewer - US
In this role, you will analyze and provide feedback on text, webpages, images, and other types of online content for leading search engines using a specialized online platform.
Â
By reviewing and rating search results for relevance and quality, you will help improve the overall search experience for millions of users around the world, including yourself.
Join our global community and put your skills to work supporting one of the world's leading search technologies.
Â
Service RatesÂ
The rate of pay is $0.2333 per completed task, with an estimated earning potential of $14 per hour. Compensation is based on tasks completed and project availability. Estimated earnings may vary depending on task volume and program requirements including, quality, and productivity, in accordance with the program's quality standards and guidelines.
Â
Please note that only one member per household may participate in this program. If it is identified that more than one person from the same household is participating in the TELUS Digital Rating Program, all associated accounts may be removed from the program.
Â
Independent Contractor Relationship
This opportunity is offered on an independent contractor basis. Contributors have the flexibility to choose when and how much they work, subject to project availability and quality requirements. This opportunity is not and should not be construed as creating an employment relationship with TELUS Digital.
Â
Qualification Process
No previous professional experience is required to apply for this opportunity. However, participation in this project requires meeting the basic requirements and successfully completing the qualification process.
Â
Basic Requirements
⢠Excellent written and verbal communication skills in English
⢠Must have resided in the United States for the past three consecutive years
⢠Familiarity with current and historical business, media, sports, news, social media, and cultural affairs in the United States
⢠Active use of Gmail and social media platforms
⢠Experience using web browsers to navigate and interact with a variety of online content
⢠Daily access to a reliable broadband internet connection
⢠Access to a smartphone (Android 5.0 or higher, or iOS 14 or higher)
⢠Access to a personal computer
Â
Assessment
You must complete an English language assessment, pass an open book qualification exam, and complete an ID verification in order to be successful for this role. These assessments will determine your suitability for the position.
To apply: https://weworkremotely.com/remote-jobs/telus-digital-content-reviewer-us-5
Expert practitioner writes, solves, and grades training tasks for AI models in your field of expertise, providing structured feedback on model outputs.
Headquarters: Atlanta, AZ.
URL: https://trainbrain.space/
What if the work you already do every day â shipping code, closing books, reading charts â became the training signal that teaches frontier AI how to think in your field?
TrainBrain builds expert-authored training data and evaluations for AI labs. We're hiring practising professionals to write, solve, and grade the tasks advanced models are measured against. When a model returns a broken function, a wrong reconciliation, or an unsafe clinical suggestion, it's because no expert caught it during training. You'd be that expert.
This isn't your day job in a new wrapper. There are no clients, no month-end close, no tickets, no on-call. You work in writing, on your own schedule, on problems chosen to sit right at the edge of what current models can handle.
No AI background required â we train you on the tooling.
Track 1 â Software Engineering
Write, debug, and optimise code across languages and problem domains. Review AI-generated solutions for correctness, security, and performance, and design edge cases that expose model weaknesses. Provide structured feedback explaining why code works â or doesn't â and compare competing solutions on engineering quality.
You'll need:
Nice to have:
Track 2 â Finance & Accounting
Author reconciliation scenarios: payment-to-invoice matching, billing versus recognised revenue, margin and variance analysis. Build the underlying data and source documents so each scenario is realistic, write rubrics specifying exact figures and required reasoning steps, then solve every task yourself to confirm the numbers hold.
You'll need:
Nice to have:
Track 3 â Medical & Clinical
Challenge models on differential diagnosis, drug interactions, treatment protocols, pathophysiology, and evidence appraisal. Write clinical vignettes with defensible, sourced answers, verify model outputs against current evidence, and document precisely where reasoning breaks down.
You'll need:
Nice to have:
What the work demands
Whatever your track, four things matter more than anything else.
Show your work. Making your reasoning explicit matters as much as the answer itself.
Be verifiable. Every task you write needs a correct answer someone else can confirm.
Iterate. You'll refine tasks until difficulty and gradability are both right.
Flag uncertainty. Saying "I'm not sure, and here's why" is a feature, not a failure.
You'll also need excellent written English, sharp attention to detail, and a secure computer with reliable internet.
Terms
This is a 1099 independent contractor position, not W-2 employment. You're responsible for your own taxes, and company-sponsored benefits such as health insurance, PTO, and retirement contributions don't apply. Contractors outside the US are engaged under the equivalent local arrangement.
Pay: $45â$100/hr, set by domain, seniority, and assessment performance. Specialised and hard-to-source expertise sits at the top of the band, and you're paid on a regular cadence.
Schedule: Choose your own projects, hours, and volume. Scale up in quiet weeks, scale down when your day job gets busy.
Growth: Strong contributors are invited into higher-rate specialist projects and review roles, with ongoing work as new engagements launch.
Provides Portuguese language feedback and annotations to train multimodal AI models.
Ex-MBB consultants create structured learning environments and tasks to train AI models on real-world consulting problem-solving and business reasoning.
Toloka AI supports frontier model post-training by building domain-specific reinforcement learning environments, tasks, and evaluation frameworks designed by real practitioners.
Mindrift, powered by Toloka â a leading enterprise AI and machine learning data partner since 2014 â connects top domain experts with cutting-edge AI initiatives. Backed by Tolokaâs deep expertise in scalable data generation, crowd technology, and applied ML systems, Mindrift enables experts to shape how next-generation generative models learn, reason, and perform.
We are launching a Management Consulting domain focused on translating real-world consulting engagements into structured learning environments for advanced AI systems. To do this credibly, we are assembling a team of strategy consultants from top-tier firms who can convert authentic project experience into end-to-end examples â from problem structuring and work planning to analysis, synthesis, and client-ready recommendations.
You will join a growing team of consultants from leading strategy firms shaping how AI learns high-level business reasoning.
Important: This role is exclusively for consultants with direct experience at a top-tier strategy consulting firm. If you do not have hands-on project experience at one of the firms listed below, please do not apply. This requirement ensures the domain is built by practitioners trained to the highest standards of structured problem-solving and client delivery.
Eligible firms: McKinsey & Company, Boston Consulting Group (BCG), Bain & Company, Oliver Wyman, Roland Berger, Monitor Deloitte (Deloitte S&C), EY-Parthenon, Kearney, and Strategy& (PwC).
Consultants with 3+ years of experience at one of the firms listed above, with hands-on project experience in:
No deep technical background is required â we will onboard you on the lightweight tools involved.
This is a remote, project-based, individual-contributor role focused on analytical design and evaluation.
On this project, contributors can earn up to $60 per hour equivalent, depending on their level and pace of contribution.
Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.
For this project, tasks are estimated to require around 25-30 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.
Ex-MBB consultant creates realistic consulting project scenarios and structured learning tasks to train AI systems on business reasoning and problem-solving.
Toloka AI supports frontier model post-training by building domain-specific reinforcement learning environments, tasks, and evaluation frameworks designed by real practitioners.
Mindrift, powered by Toloka â a leading enterprise AI and machine learning data partner since 2014 â connects top domain experts with cutting-edge AI initiatives. Backed by Tolokaâs deep expertise in scalable data generation, crowd technology, and applied ML systems, Mindrift enables experts to shape how next-generation generative models learn, reason, and perform.
We are launching a Management Consulting domain focused on translating real-world consulting engagements into structured learning environments for advanced AI systems. To do this credibly, we are assembling a team of strategy consultants from top-tier firms who can convert authentic project experience into end-to-end examples â from problem structuring and work planning to analysis, synthesis, and client-ready recommendations.
You will join a growing team of consultants from leading strategy firms shaping how AI learns high-level business reasoning.
Important: This role is exclusively for consultants with direct experience at a top-tier strategy consulting firm. If you do not have hands-on project experience at one of the firms listed below, please do not apply. This requirement ensures the domain is built by practitioners trained to the highest standards of structured problem-solving and client delivery.
Eligible firms: McKinsey & Company, Boston Consulting Group (BCG), Bain & Company, Oliver Wyman, Roland Berger, Monitor Deloitte (Deloitte S&C), EY-Parthenon, Kearney, and Strategy& (PwC).
Consultants with 3+ years of experience at one of the firms listed above, with hands-on project experience in:
No deep technical background is required â we will onboard you on the lightweight tools involved.
This is a remote, project-based, individual-contributor role focused on analytical design and evaluation.
On this project, contributors can earn up to $60 per hour equivalent, depending on their level and pace of contribution.
Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.
For this project, tasks are estimated to require around 25-30 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.
Ex-MBB consultant creates realistic consulting project scenarios and structured tasks to train AI models on business reasoning and problem-solving.
Toloka AI supports frontier model post-training by building domain-specific reinforcement learning environments, tasks, and evaluation frameworks designed by real practitioners.
Mindrift, powered by Toloka â a leading enterprise AI and machine learning data partner since 2014 â connects top domain experts with cutting-edge AI initiatives. Backed by Tolokaâs deep expertise in scalable data generation, crowd technology, and applied ML systems, Mindrift enables experts to shape how next-generation generative models learn, reason, and perform.
We are launching a Management Consulting domain focused on translating real-world consulting engagements into structured learning environments for advanced AI systems. To do this credibly, we are assembling a team of strategy consultants from top-tier firms who can convert authentic project experience into end-to-end examples â from problem structuring and work planning to analysis, synthesis, and client-ready recommendations.
You will join a growing team of consultants from leading strategy firms shaping how AI learns high-level business reasoning.
Important: This role is exclusively for consultants with direct experience at a top-tier strategy consulting firm. If you do not have hands-on project experience at one of the firms listed below, please do not apply. This requirement ensures the domain is built by practitioners trained to the highest standards of structured problem-solving and client delivery.
Eligible firms: McKinsey & Company, Boston Consulting Group (BCG), Bain & Company, Oliver Wyman, Roland Berger, Monitor Deloitte (Deloitte S&C), EY-Parthenon, Kearney, and Strategy& (PwC).
Consultants with 3+ years of experience at one of the firms listed above, with hands-on project experience in:
No deep technical background is required â we will onboard you on the lightweight tools involved.
This is a remote, project-based, individual-contributor role focused on analytical design and evaluation.
On this project, contributors can earn up to $60 per hour equivalent, depending on their level and pace of contribution.
Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.
For this project, tasks are estimated to require around 25-30 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.
Transcribes audio in Slovenian to create training data for AI model development and evaluation.
Transcribes audio in Egyptian Arabic to create high-quality training data for AI models.
Transcribes audio in Brazilian Portuguese to create high-quality training data for AI models.
Test and provide feedback on AI notification-sorting workflows on Google Pixel devices while using apps naturally and completing daily check-in forms.
Hindi AI Product Tester
Active Project Status: This project is now live, and we are urgently onboarding participants to begin immediately. Because we are moving quickly, we kindly ask that you keep a close eye on your communication (email and messages) after applying. It is essential that you respond promptly so we can get you cleared and ready to start without delay.
Location: 100% Remote, global
Device: Applicants must own a qualifying Google Pixel device for the duration of the project
Compensation:USD $140.00 upon successful completion of the 14 qualified submissions (each submission should take 10â15 minutes of activity per day)
Project Requirements
The Project
We are seeking detail-oriented and motivated speakers to participate in an upcoming notification sorting project.
Participants will use a Google Pixel 9 or newer device throughout the project period while interacting naturally with apps that generate notifications. Participants will complete daily project tasks involving notification organization/sorting workflows and provide related feedback/documentation as instructed.
Participants will also complete a short daily feedback/check-in form throughout the duration of the pilot to assist with participation tracking and troubleshooting support.
This project is part of an ongoing effort to improve emerging mobile AI and notification-management technologies for international users.
Workload
Duties
Engagement Requirements
All participants must have, or be willing to create, an Upwork account. The project will be managed via the Upwork platform.
While participating in this project, adherence to the confidentiality terms outlined in the Upwork User Agreement is required. Any information accessed or received related to this project is confidential and may not be shared or disclosed to third parties.
Participants must:
Use their own qualifying Google Pixel device.
Qualifying Device Examples
Google Pixel 9
Google Pixel 9 Pro
Google Pixel 10
Google Pixel Fold (newer generation)
Please note:Google Pixel 9a and Pixel 10a devices are not compatible with project requirements.
About Us:
As a global data company, Productive Playhouse âPPHâ, is pioneering our approach to language and data services, while incorporating their roots as a production company. Originally creating content to support childrenâs language acquisition, our commitment to excellence, forward-thinking strategies, and world-wide cultural experience has proven key for delivering exceptional service.
Originally founded as an educational production company, Productive Playhouse made a mark with our award winning childrenâs series, which taught fundamental subjects through engaging and effective programming. This early success paved the way for our evolution into a comprehensive data services provider.
Since 2011, Productive Playhouse has expanded rapidly to offer an extensive suite of data services. Our current offerings include transcription, translation, linguistic analysis, rating, systems testing, localization, field and studio recording, language skill verification, and specialized data handling with a focus on sensitivity and diversity. Our commitment to innovation means we continually enhance our service portfolio to meet the evolving needs of our clients.
At Productive Playhouse, we are proud of our reputation for addressing complex challenges with agility and delivering premium, secure data solutions across diverse environments. Our dynamic team is dedicated to maintaining the highest standards and ensuring exceptional service every time.
Disclaimer:
All engagements are contingent upon successful completion of a background screening. Productive Playhouse is committed to diversity and inclusion across our global contractor network. We welcome applicants of all backgrounds and abilities.
Mentors and teaches developers on building agentic AI systems as an independent contractor across multiple regions.
Applies garment manufacturing QC expertise to train AI models by providing feedback and labeling data for quality control automation systems.
Evaluates and rates music outputs to provide feedback for AI model training and improvement.