AI Safety Policy
Safety-First Design
Every Leuks model is developed with safety as a first-class design requirement—not a post-hoc addition. Our safety commitments are embedded into the training process, evaluation pipeline, and deployment infrastructure. We believe that safety and capability are complementary, not competing, goals.
Red Teaming Program
Before any major model release, we conduct structured red teaming exercises involving internal safety researchers and—where possible—external domain experts. Red teamers are tasked with systematically attempting to elicit unsafe, biased, or otherwise harmful outputs across a wide range of scenarios.
Red teaming categories include:
- Harmful content generation (violence, CSAM, self-harm)
- Disinformation and manipulation
- Technical uplift for dangerous activities
- Bias and demographic harm
- Jailbreak and prompt injection resistance
- Privacy violations and data extraction
Findings from red teaming are used to improve model safety before public release.
Alignment Approach
Our alignment work focuses on ensuring that Leuks models reliably follow instructions, remain helpful, and avoid harmful behaviors—even when faced with ambiguous or adversarial inputs. We use Reinforcement Learning from Human Feedback (RLHF), Constitutional AI-inspired techniques, and dedicated safety fine-tuning to align model behavior with human values and our policies.
We acknowledge that alignment is an unsolved research problem and commit to ongoing investment in alignment research as the field evolves.
Content Moderation
Runtime content moderation complements model-level safety. Our moderation pipeline includes:
- Automated classifiers that flag potentially harmful inputs and outputs
- Real-time blocking of content in the highest-severity categories
- Human review queues for ambiguous flags
- Feedback loops that route moderation decisions back into model improvement
Incident Response
When a safety incident is identified—whether by internal monitoring, user reports, or external researchers—we follow a structured response process:
- Triage: Assess severity and scope within 2 hours of identification
- Containment: Apply immediate mitigations (rule updates, access restrictions) as needed
- Investigation: Root cause analysis within 5 business days
- Remediation: Model or policy updates with appropriate testing
- Post-mortem: Document learnings and update processes to prevent recurrence
Significant safety incidents that affect users at scale will be disclosed in our Transparency Report.
Reporting Safety Issues
We welcome reports of safety issues from researchers, users, and the public. To report a safety concern:
- AI safety issues: [email protected]
- Security vulnerabilities: [email protected]
- Urgent CSAM or imminent harm: Contact us immediately and report to relevant national authorities
We commit to acknowledging all safety reports within 3 business days.
Research Commitments
Inferent is committed to advancing the field of AI safety through:
- Publishing research findings where possible without compromising security
- Collaborating with academic institutions and safety organizations
- Participating in industry-wide safety initiatives and standards development
- Supporting external audits of our safety practices