AI Response Labeler / Annotator at Blueprint Technologies
Worldwide
$2.2k - $2.5k
<div class="content-intro"><p><span style="color: rgb(57, 116, 216);"><strong><span style="font-size: 24pt;">About Blueprint</span></strong></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">Blueprint is a technology solutions firm headquartered in Bellevue, Washington, with teams across the United States. We help organizations turn complex challenges into meaningful outcomes by connecting strategy and execution across AI, cloud, data, product development, and emerging technology.</span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">Our culture is built by people who care deeply about doing exceptional work. We set high standards, take ownership, and continually challenge ourselves and one another to be better. We work hard, support each other, and take genuine pride in what we deliver for our clients, partners, and teams.</span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">At Blueprint, you’ll work alongside talented people with different experiences, expertise, and perspectives. You’ll have opportunities to take on meaningful challenges, expand your skills, and see the impact of what you build.</span></p> <p><strong><span style="font-size: 18pt; color: rgb(57, 116, 216);">Bring your perspective. Raise the standard. Build what matters.</span></strong></p></div><p><span style="color: rgb(57, 116, 216);"><strong><span style="font-size: 24pt;">About the Role</span></strong></span></p> <p class="isSelectedEnd"></p> <p class="isSelectedEnd"><span style="font-size: 14pt; color: rgb(32, 33, 36);">We’re looking for an <strong>English-language AI Response Labeler / Annotator</strong> to evaluate the quality of AI-generated responses. This role calls for strong English comprehension, analytical judgment, and the ability to apply detailed guidelines consistently across a high volume of work.</span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">You’ll compare responses generated by different AI models and determine which one better meets a user’s needs. You’ll consider factual accuracy, reasoning, relevance, completeness, instruction following, safety, clarity, tone, and overall usefulness. The prompts, responses, annotation guidelines, training, and written evaluation work for this role are in English.</span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="color: rgb(57, 116, 216);"><strong><span style="font-size: 24pt;">What You'll Do</span></strong></span></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"></span></p> <ul> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Perform side-by-side comparisons of AI-generated responses and select the stronger response using established evaluation criteria.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Assess responses for factual accuracy, relevance, completeness, reasoning, instruction following, clarity, safety, tone, and usefulness.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Evaluate varied tasks, including questions and answers, web-search results, file-based and image-based responses, content generation, and single-turn or multi-turn conversations.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Identify meaningful differences between responses, such as unsupported claims, missed instructions, weak reasoning, and incomplete answers.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Apply detailed, scenario-specific guidelines and make sound decisions when an example does not provide an obvious answer.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Write concise, evidence-based explanations for your decisions when required.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Meet established productivity expectations while maintaining accuracy and consistent judgment.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Participate in training, guided practice, calibration, qualification reviews, and ongoing quality reviews.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Incorporate feedback as evaluation guidelines and quality standards evolve.</span></li> </ul> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="color: rgb(57, 116, 216);"><strong><span style="font-size: 24pt;">What You'll Bring</span></strong></span></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"></span></p> <ul> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Excellent written English comprehension and communication skills, including the ability to read complex prompts and guidelines and explain evaluation decisions clearly.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Strong critical-thinking skills and the ability to assess content across a wide range of topics.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Sound judgment when evaluating factuality, reasoning, user intent, and subtle differences in response quality.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Excellent attention to detail and the ability to apply structured criteria consistently.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Comfort with repetitive, focused work and a high volume of evaluations.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Ability to work independently, respond to feedback, and stay aligned with shared quality standards.</span></li> </ul> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"><strong><span style="color: rgb(57, 116, 216);"><span style="font-size: 24pt;">Preferred Qualifications</span></span></strong></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"></span></p> <ul> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Experience evaluating, ranking, or comparing AI-generated responses, particularly through side-by-side evaluation.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Experience with data annotation, content quality assessment, search relevance evaluation, or model-quality review.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Experience working with detailed rubrics, annotation guidelines, or quality benchmarks.</span></li> <li style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Experience writing clear rationales that support evaluation decisions.</span></li> </ul> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="color: rgb(57, 116, 216);"><strong><span style="font-size: 24pt;">Work Pace and Productivity Expectations</span></strong></span></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">This is a highly structured and repetitive role that involves completing similar evaluation tasks throughout the workday. Candidates should be comfortable maintaining focus, accuracy, and consistent judgment while reviewing a high volume of AI-generated content.</span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">Most evaluation tasks are expected to take approximately 15 minutes, and employees are generally expected to complete a minimum of 25 tasks per day. Some tasks may take more or less time depending on their complexity.</span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">Success in this role requires balancing productivity with quality. Employees must meet established daily expectations while carefully applying annotation guidelines and providing accurate, well-supported evaluation decisions.</span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="color: rgb(57, 116, 216);"><strong><span style="font-size: 24pt;">Training and Qualification</span></strong></span></span></p> <p class="isSelectedEnd"><span style="font-size: 14pt; color: rgb(32, 33, 36);">All new hires must successfully complete a structured onboarding and qualification program before beginning production work.</span></p> <p class="isSelectedEnd"><span style="font-size: 14pt; color: rgb(32, 33, 36);">The program includes training sessions, guided practice exercises, calibration against established quality benchmarks, and a formal qualification review.</span></p> <p class="isSelectedEnd"><span style="font-size: 14pt; color: rgb(32, 33, 36);">Training is intended to establish consistent evaluation judgment across the team. Language fluency alone will not be sufficient to qualify. Employees must also demonstrate the ability to evaluate broader response quality, follow detailed annotation guidelines, explain their decisions, and complete work within the expected timeframe.</span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">Employees will continue to receive feedback, quality reviews, and calibration support after entering production.</span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"><span style="color: rgb(57, 116, 216);"><strong><span style="font-size: 24pt;">Compensation</span></strong></span></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">The estimated compensation range is <strong>USD $2,200–$2,500 per month ($26,400–$30,000 annually)</strong>. Actual compensation will depend on the hiring location, experience, skills, and internal equity. Compensation may be paid in local currency through the applicable local Professional Employer Organization (PEO) or Employer of Record (EOR) partner.</span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);"><strong><span style="color: rgb(57, 116, 216);"><span style="font-size: 24pt;">Location and Employment Structure</span></span></strong></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">This is a remote role open to candidates in <strong>Latin American countries</strong>. Employment will be arranged through a local PEO or EOR partner, as applicable. Payroll, statutory benefits, and employment terms will follow the requirements of the candidate’s hiring country and the terms of their employment.</span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">During the approximately 30-day training and qualification period, employees must work from 9:00 a.m. to 5:00 p.m. Pacific Time. After successfully completing training, employees may work standard business hours within their local time zone.</span></p><div class="content-conclusion"><p><span style="color: rgb(57, 116, 216);"><strong><span style="font-size: 24pt;">Benefits</span></strong></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">Blueprint believes that healthy, supported employees do their best work. Eligible employees have access to a comprehensive benefits package that may include:</span></p> <ul> <li><span style="font-size: 14pt; color: rgb(32, 33, 36);">Medical, dental, and vision coverage</span></li> <li><span style="font-size: 14pt; color: rgb(32, 33, 36);">Flexible Spending Account (FSA)</span></li> <li><span style="font-size: 14pt; color: rgb(32, 33, 36);">401(k) retirement plan</span></li> <li><span style="font-size: 14pt; color: rgb(32, 33, 36);">Competitive paid time off</span></li> <li><span style="font-size: 14pt; color: rgb(32, 33, 36);">Parental leave</span></li> <li><span style="font-size: 14pt; color: rgb(32, 33, 36);">Professional growth and development opportunities</span></li> </ul> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">Benefits and eligibility may vary based on role, employment status, and location.</span></p> <p><span style="color: rgb(57, 116, 216);"><strong><span style="font-size: 24pt;">Equal Employment Opportunity</span></strong></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">Blueprint Technologies, LLC is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, pregnancy, childbirth or related medical conditions, sexual orientation, gender identity or expression, national origin, ancestry, age, disability, genetic information, marital or familial status, military or veteran status, citizenship status, or any other characteristic protected by applicable law.</span></p> <p><span style="color: rgb(57, 116, 216);"><strong><span style="font-size: 24pt;">Applicant Accommodations</span></strong></span></p> <p><span style="font-size: 14pt; color: rgb(32, 33, 36);">If you need a reasonable accommodation to participate in any part of the application or interview process, please contact <a href="mailto:recruiting@bpcs.com">recruiting@bpcs.com.</a></span></p></div>
Apply Now