Capability and factuality
Whether the model gives a factually correct answer and retains useful task capability.
Open Research Project 01
Research groundwork underwayAmharic phaseThis project measures how Amharic language adaptation and agricultural fine-tuning change model capability, refusal, uncertainty and locally relevant advice.
Underlying model-adaptation research is already underway. The funded evaluation phase begins after the minimum funding threshold or equivalent project-specific institutional funding is confirmed.
External funding toward this phase
$0
of $27,500 full target
Groundwork already funded and underway. Addis AI has funded the underlying model research and infrastructure; external funding expands this evaluation phase.
How do continued pretraining, broad instruction tuning and agricultural specialization change the alignment and usefulness of an open model in Amharic?
Teams are adapting open models to languages and domains with limited evaluation of what changes at each stage. Agriculture makes the problem concrete because local conditions, missing context and confident errors can change the value of an answer.
What already exists
Existing research
Addis AI has previously carried out vocabulary extension, continued pretraining and supervised fine-tuning for Amharic. That work establishes the underlying adaptation pipeline and informs this study, but those earlier checkpoints are not treated as the controlled stage sequence for the research described here.
The new study will use a separately frozen experimental configuration so that changes introduced during language adaptation, instruction tuning and domain specialization can be evaluated consistently.
The links below show Addis AI's broader open-source and benchmark track record. They are not presented as the controlled stage sequence for this study.
Stage comparison
Start from an instruction-tuned open model with an existing safety and instruction-following baseline. The exact checkpoint and configuration will be selected and frozen in the public protocol before the main evaluation begins.
Apply the frozen Amharic language-adaptation configuration, then measure capability, factuality, language, local relevance, safety and alignment behavior again. The same agriculture suite is used so changes can be compared with the earlier checkpoints.
Teach general Amharic instruction following, then repeat the same evaluation suite. The same agriculture suite is used so changes can be compared with the earlier checkpoints.
Apply agriculture or health supervised fine-tuning, then compare the change from every earlier stage. The same agriculture suite is used so changes can be compared with the earlier checkpoints.
The protocol will be published before the main stage-comparison evaluation begins.
The protocol will state whether tokenizer changes are treated as part of the language-adaptation stage, separately controlled, or held fixed for the stage comparison.
Research questions
Evaluation and metrics
Whether the model gives a factually correct answer and retains useful task capability.
Whether it understands the Amharic input and produces an answer appropriate to the relevant local context.
Whether following the response could create meaningful harm.
Whether adaptation changes behaviors such as harmful compliance, inappropriate refusal, unsupported certainty, escalation or other safety-relevant behavior.
Scenario counts, reviewer redundancy, slice definitions and the statistical protocol will be fixed before the main evaluation. The public protocol will state exclusions and any later changes.
Reference standards
Reference selection and local-context decisions will be documented in the evaluation protocol.
Human and domain evaluation
Native Amharic evaluators test language quality and practical usefulness. Agronomists review factuality, missing context, severity and locally inappropriate recommendations.
Execution plan
Relative targets begin at project start. Actual dates will be added to the research log after funding is confirmed.
Mitigation experiments
A mitigation is tested only after the stage comparison identifies a failure or unwanted change. The affected metrics are then measured again.
Full scope
Mix a controlled sample from the earlier training distribution into adaptation and test whether it limits drift.
Full scope
Add selected safety and alignment examples at the stage where evaluation finds a change, then re-run the suite.
Minimum and full scope
Remove or reweight domain examples linked to factual, contextual or safety failures and compare the result.
Conditional
Test LoRA or a smaller adaptation capacity only if the first comparison suggests that parameter change is part of the failure.
Stretch
Use a small DPO or preference-tuning experiment only if reviewer data and the remaining budget support it.
Research log
Targets remain relative until funding is confirmed. When work starts, the log can add planned dates, actual completion dates, artifacts and delay or change notes without replacing the original target window.
Protocol, source inventory, reviewer onboarding and evaluation-set construction.
Baseline, CPT and broad-SFT evaluation.
Agriculture-SFT evaluation, native review and agronomist review.
Failure analysis, initial results and minimum-scope public release.
Mitigation experiments and expanded contextual evaluation.
Repeat evaluation, analysis and full public release.
Funding scope
The minimum funds a smaller but scientifically useful study with expert review. The expansion increases coverage, redundancy, mitigation testing, repeatability and release depth. Each work package appears once in the allocation below.
Addis AI funds the core engineering team and existing infrastructure. External funding primarily supports human evaluation, domain expertise, field research, additional experiments and open release.
Minimum useful scope
$14,000
Full-scope expansion
$13,500
Full target
$27,500
Minimum useful scope
$14,000
A smaller but scientifically useful study with expert review, a four-stage comparison and a reproducible public release.
Full research target
$27,500
The additional funding increases evaluation coverage, reviewer redundancy, contextual evaluation, mitigation testing, repeatability, statistical confidence and reproducibility.
What the additional $13,500 adds
Open deliverables
Beyond Addis AI
The protocol and stage comparisons can be reused by teams running continued pretraining or supervised fine-tuning for other low-resource languages. There is little public evidence showing how alignment changes during this process.
Risks and limitations
Research updates
We publish progress, changes in scope, negative results and delays as they happen.
Milestone · Protocol and source review
The public scope, stage comparison, initial budget and planned outputs are available for review.
Findings: No project results are claimed at this stage.
Next step: Freeze the scenario, reviewer and statistical protocols before the main evaluation.
Funding
Funding totals include only cleared contributions and confirmed project-specific grants or sponsorships. Read the funding policy before contributing.
Addis AI funds the core engineering team and existing infrastructure. External funding primarily supports human evaluation, domain expertise, field research, additional experiments and open release.
Institutional funders can support a defined project work package through a separate project-specific agreement.
External funding toward this phase
$0
of $27,500 full target
Groundwork already funded and underway. Addis AI has funded the underlying model research and infrastructure; external funding expands this evaluation phase.
Funding totals include only cleared contributions and confirmed project-specific grants or sponsorships.
Funding above the target
Funding above the target first expands Amharic sampling and independent re-review. Work in another language requires a separately published scope.
Institutional and research contact