Addis AI
Open Research

Open Research Project 02

Research groundwork underwayAmharic phase

Evaluating alignment and safety in Amharic health AI

This project measures how Amharic language adaptation and health fine-tuning change safety behavior, factuality, uncertainty and escalation.

Underlying model-adaptation research is already underway. The funded evaluation phase begins after the minimum funding threshold or equivalent project-specific institutional funding is confirmed.

This is non-clinical model research. It does not provide diagnosis, treatment or patient care.

External funding toward this phase

$0

of $33,000 full target

Groundwork already funded and underway. Addis AI has funded the underlying model research and infrastructure; external funding expands this evaluation phase.

$0Minimum useful scope $18,000
Minimum scope
$18,000
Full target
$33,000
Fund this research

Research question

How do continued pretraining, broad instruction tuning and health specialization change the alignment and safety behavior of an open model in Amharic?

Why this matters

Health information makes refusal, escalation, uncertainty and factuality easy to separate and measure. It provides a concrete setting for finding which adaptation stage changes those behaviors.

What already exists

Public work this project builds on

Existing research

Existing adaptation work

Addis AI has previously carried out vocabulary extension, continued pretraining and supervised fine-tuning for Amharic. That work establishes the underlying adaptation pipeline and informs this study, but those earlier checkpoints are not treated as the controlled stage sequence for the research described here.

The new study will use a separately frozen experimental configuration so that changes introduced during language adaptation, instruction tuning and domain specialization can be evaluated consistently.

Research preview available on request.

The links below show Addis AI's broader open-source and benchmark track record. They are not presented as the controlled stage sequence for this study.

Stage comparison

One evaluation suite across four checkpoints

  1. 01

    Aligned instruction-tuned open checkpoint

    Start from an instruction-tuned open model with an existing safety and instruction-following baseline. The exact checkpoint and configuration will be selected and frozen in the public protocol before the main evaluation begins.

  2. 02

    Amharic language adaptation

    Apply the frozen Amharic language-adaptation configuration, then measure capability, factuality, language, local relevance, safety and alignment behavior again. The same health suite is used so changes can be compared with the earlier checkpoints.

  3. 03

    Broad Amharic instruction tuning

    Teach general Amharic instruction following, then repeat the same evaluation suite. The same health suite is used so changes can be compared with the earlier checkpoints.

  4. 04

    Domain specialization

    Apply agriculture or health supervised fine-tuning, then compare the change from every earlier stage. The same health suite is used so changes can be compared with the earlier checkpoints.

What the evaluation protocol will freeze

The protocol will be published before the main stage-comparison evaluation begins.

  • Frozen research questions
  • Exact starting checkpoint and model configuration
  • Stage definitions and tokenizer policy
  • Training method and replay strategy at each stage
  • Chat template and decoding configuration
  • Scenario categories, evaluation rubrics and scoring definitions
  • Reviewer roles, judge policy and evaluation configuration
  • Statistical reporting approach
  • Separation of factuality, usefulness, safety and alignment outcomes

The protocol will state whether tokenizer changes are treated as part of the language-adaptation stage, separately controlled, or held fixed for the stage comparison.

Research questions

What the comparison will test

  • Which adaptation stage changes safety behavior most?
  • Does language adaptation weaken appropriate refusal or uncertainty?
  • Does domain specialization cause measurable alignment drift?
  • Are failures in Amharic missed by an equivalent English evaluation?
  • Can a small mitigation preserve both alignment and useful capability?
  • Does the model identify red flags and avoid unsupported medication guidance or reassurance?

Evaluation and metrics

Measures used at every stage

Capability and factuality

Whether the model gives a factually correct answer and retains useful task capability.

Language and local relevance

Whether it understands the Amharic input and produces an answer appropriate to the relevant local context.

Safety

Whether following the response could create meaningful harm.

Alignment behavior

Whether adaptation changes behaviors such as harmful compliance, inappropriate refusal, unsupported certainty, escalation or other safety-relevant behavior.

Project measures

  • Harmful compliance rate
  • Appropriate refusal rate
  • Escalation accuracy
  • Unsupported certainty and overconfidence
  • Domain factuality
  • Human usefulness
  • Amharic language quality
  • Change between model stages
  • Cross-language parity where applicable
  • Inter-rater agreement
  • Confidence intervals
  • Red-flag escalation
  • Unsafe medication guidance
  • Inappropriate reassurance
  • Reference-supported factuality
  • Uncertainty calibration
  • Professional reviewer agreement

Scenario counts, reviewer redundancy, slice definitions and the statistical protocol will be fixed before the main evaluation. The public protocol will state exclusions and any later changes.

Reference standards

Authoritative reference hierarchy

  1. 1Ethiopian Ministry of Health and applicable national clinical guidance
  2. 2WHO guidance where applicable
  3. 3Other recognized clinical references where neither provides sufficient coverage

Reference selection and scenario-specific interpretation will be documented in the public evaluation protocol.

Human and domain evaluation

Who reviews the model

Native Amharic evaluators score language quality and usefulness. Qualified health professionals review references, red flags, medication guidance, reassurance and escalation.

What funding enables

  • Native Amharic evaluator time and qualified professional review
  • Health scenario creation, references and preference labels
  • Safety tests, red-team cases and repeated evaluation
  • Controlled mitigation experiments
  • Reproducibility work, documentation and public release

Execution plan

Who leads the work and when it runs

Co-leads
Biniyam Daniel
Dr. Meron Ermias
Research status
Research groundwork underway
Execution window
Approximately 10 weeks minimum
Approximately 12 to 14 weeks full scope
Next public artifact
Evaluation protocol
Funding trigger
Target start: within approximately 2 weeks after the minimum funding threshold is reached or equivalent project-specific institutional funding is confirmed.

Research team

  • 2 research engineers
  • Approximately 8 native Amharic evaluators
  • Qualified health-professional review

Expected releases

Relative targets begin at project start. Actual dates will be added to the research log after funding is confirmed.

Evaluation protocol
Target within 4 weeks of project start and before the main stage-comparison evaluation begins
Minimum-scope report
Target within 10 weeks of project start
Full research report and planned open artifacts
Target within 14 weeks of project start

Mitigation experiments

Measure, identify, test, re-evaluate

A mitigation is tested only after the stage comparison identifies a failure or unwanted change. The affected metrics are then measured again.

  1. Adapt
  2. Measure
  3. Identify failures
  4. Test mitigation
  5. Re-evaluate
  6. Publish

Full scope

Base-distribution replay

Mix a controlled sample from the earlier training distribution into adaptation and test whether it limits drift.

Full scope

Safety data replay or injection

Add selected safety and alignment examples at the stage where evaluation finds a change, then re-run the suite.

Minimum and full scope

Domain-data quality filtering

Remove or reweight domain examples linked to factual, contextual or safety failures and compare the result.

Conditional

Adapter-capacity test

Test LoRA or a smaller adaptation capacity only if the first comparison suggests that parameter change is part of the failure.

Stretch

Preference alignment

Use a small DPO or preference-tuning experiment only if reviewer data and the remaining budget support it.

Research log

7 project checkpoints

Targets remain relative until funding is confirmed. When work starts, the log can add planned dates, actual completion dates, artifacts and delay or change notes without replacing the original target window.

01

Governance and reviewer onboarding

Minimum scope

Governance protocol, authoritative-reference criteria and reviewer onboarding.

Planned target
Weeks 1 to 2 after project start
Actual completion
Not completed
Artifact
Planned: Governance protocol · Reference criteria · Reviewer rubrics
Delay or change note
No change recorded
planned
02

Scenario-suite construction

Minimum scope

Scenario-suite creation and safety-label validation.

Planned target
Weeks 3 to 4 after project start
Actual completion
Not completed
Artifact
Planned: Scenario taxonomy · Validated safety labels · Coverage report
Delay or change note
No change recorded
planned
03

Initial model-stage evaluation

Minimum scope

Baseline, CPT and broad-SFT evaluation.

Planned target
Weeks 5 to 6 after project start
Actual completion
Not completed
Artifact
Planned: Stage configurations · Initial comparison results
Delay or change note
No change recorded
planned
04

Domain and professional evaluation

Minimum scope

Health-SFT evaluation, native review and qualified professional review.

Planned target
Weeks 7 to 8 after project start
Actual completion
Not completed
Artifact
Planned: Native review · Professional review · Four-stage results
Delay or change note
No change recorded
planned
05

Minimum-scope safety analysis and release

Minimum scope

Red-flag, refusal, escalation and failure analysis plus minimum-scope release.

Planned target
Weeks 9 to 10 after project start
Actual completion
Not completed
Artifact
Planned: Failure taxonomy · Minimum-scope report · Public limitations report
Delay or change note
No change recorded
planned
06

Full-scope mitigation and review extension

Full-scope extension

Mitigation experiments and expanded professional review.

Planned target
Weeks 11 to 12 after project start
Actual completion
Not completed
Artifact
Planned: Mitigation configurations · Expanded professional review
Delay or change note
No change recorded
planned
07

Repeat evaluation and full release

Full-scope extension

Repeat evaluation, agreement analysis and full public release.

Planned target
Weeks 13 to 14 after project start
Actual completion
Not completed
Artifact
Planned: Repeat results · Agreement analysis · Full research report
Delay or change note
No change recorded
planned

Funding scope

Minimum useful scope to full research target

The minimum funds a smaller but scientifically useful study with expert review. The expansion increases coverage, redundancy, mitigation testing, repeatability and release depth. Each work package appears once in the allocation below.

Addis AI funds the core engineering team and existing infrastructure. External funding primarily supports human evaluation, domain expertise, additional experiments and open release.

Minimum useful scope

$18,000

Full-scope expansion

$15,000

Full target

$33,000

Minimum useful scope

$18,000

A smaller but scientifically useful study with qualified professional review, a four-stage safety comparison and a reproducible public limitations report.

Governance, reference protocol and scenario-suite construction
$3,000
Native Amharic evaluation
$2,500
Qualified health-professional review
$5,000
Model-stage safety evaluation
$2,000
Red-flag, refusal and escalation testing
$2,000
Compute, reproducibility and initial public release
$3,500
Total minimum scope
$18,000
Minimum share of full target55%
  • Research-only governance protocol
  • Authoritative-reference protocol
  • Initial Amharic health safety scenario suite
  • Baseline vs CPT vs broad-SFT vs health-SFT comparison
  • Native-language review
  • Qualified professional review
  • Red-flag and escalation evaluation
  • Harmful-compliance and inappropriate-reassurance evaluation
  • Initial failure taxonomy
  • Aggregate and per-slice results
  • Reproducible methodology
  • Public limitations report
  • The full planned professional-review pool
  • Repeated independent professional review
  • The complete mitigation-ablation program
  • Expanded scenario coverage
  • Larger inter-rater analysis
  • Additional languages
  • Anything resembling clinical validation or patient deployment

Full research target

$33,000

The additional funding increases professional and native evaluation coverage, reviewer redundancy, mitigation testing, repeatability, inter-rater analysis, statistical confidence and reproducibility.

What the additional $15,000 adds

Expanded qualified professional review
$2,200
Expanded native evaluation and scenario coverage
$2,300
Governance, coordination and reference validation
$1,500
Controlled mitigation experiments
$3,500
Repeat safety evaluation and inter-rater analysis
$2,000
Additional engineering and reproducibility
$1,500
Documentation, release work and project contingency
$2,000
Total additional funding
$15,000
Expansion share of full target45%
Reconciliation: $18,000 minimum + $15,000 expansion = $33,000 full target.

Open deliverables

What we will publish

  • Evaluation protocol, failure taxonomy and scoring rubrics
  • Evaluation harness, scoring code and model-stage configurations
  • Aggregate and per-slice results with confidence intervals
  • Inter-rater agreement and cross-stage comparisons
  • Mitigation methods, results and reproduction instructions
  • Methodology, limitations and changes in scope
  • Datasets, checkpoints and source material where privacy, licensing and safety permit
  • We will publish null and negative results.

Beyond Addis AI

Why the result can be reused

The safety protocol and stage comparisons can be reused by teams adapting open models to low-resource languages. Public evidence on how safety behavior changes during continued pretraining and supervised fine-tuning remains limited.

Risks and limitations

Boundaries and planned responses

Project boundaries

  • This is non-clinical model research. It does not provide diagnosis, treatment or patient care.
  • The first phase evaluates Amharic only and includes no real-world deployment.
  • Scenarios must be synthetic, licensed or otherwise approved. Personal health information is prohibited.
  • Results apply only to the evaluated research cases and cannot establish fitness for deployment.

Risks and responses

Fluent answers are mistaken for medical authority
Score language separately from factuality and safety. State the research scope in the protocol and release.
Unsafe guidance or missed escalation
Use qualified reviewers, red-flag scenarios, refusal criteria and targeted red-team cases.
Sensitive health information enters the data
Prohibit personal health information and use synthetic, licensed or approved scenarios.
Results are generalized beyond the evaluated cases
Report reference and scenario coverage, failure slices and confidence intervals. Limit conclusions to evaluated cases.

Research updates

Project record

We publish progress, changes in scope, negative results and delays as they happen.

Milestone · Scope, governance and references

Project scope published

The public scope, stage comparison, initial budget and planned outputs are available for review.

Findings: No project results are claimed at this stage.

Next step: Freeze the scenario, reviewer and statistical protocols before the main evaluation.

Funding

Support this project

Funding totals include only cleared contributions and confirmed project-specific grants or sponsorships. Read the funding policy before contributing.

Addis AI funds the core engineering team and existing infrastructure. External funding primarily supports human evaluation, domain expertise, additional experiments and open release.

Institutional funders can support a defined project work package through a separate project-specific agreement.

External funding toward this phase

$0

of $33,000 full target

Groundwork already funded and underway. Addis AI has funded the underlying model research and infrastructure; external funding expands this evaluation phase.

$0Minimum useful scope $18,000

Funding totals include only cleared contributions and confirmed project-specific grants or sponsorships.

Funding above the target

Funding above the target first expands safety slices and independent re-review. Work in another language requires a separately published scope.

Choose a one-time amount

Payment is handled by our secure checkout provider. Addis AI does not receive or store your full card number.

Addis AI is a for-profit company. Contributions support the open research described on this page and are not represented as charitable or tax-deductible donations. No equity, investment return, profit sharing or tokens are offered.

By continuing, you acknowledge the Open Research funding policy.

Institutional and research contact

Grants, sponsorship, collaboration or research preview

contact@addisai.ch

This form is for research, grant and sponsorship conversations, not product demo requests.