Text such as `[/INST]` in a prompt or response was parsed as Rich markup, which crashed the run with a MarkupError. Other bracketed text like `[bird migration]` was silently removed from the output.
* Cap ARA's neighbor count at the number of prompts
The nearest neighbors are selected from the outputs for the good and
bad prompts, which have one row per prompt. With fewer than 15 prompts
in either dataset, the optimizer could pick a neighbor count larger
than that, and torch.topk would crash with "selected index k out of
range". Bounding the search range instead of clamping the value keeps
the optimizer from wasting trials on counts that would all behave the
same.
* Render chat templates with a fixed date
Some chat templates, including gpt-oss's, put the current date in the
system prompt. That makes the residuals, and with them the modified
model, depend on the day Heretic runs, so a model could not be
reproduced on a different day. It also made the gpt-oss test's hash
change daily once ARA started modifying the model.
Prompts used for computing residuals and for evaluation are now
rendered with a fixed date. Interactive chat keeps the real one.
* Prevent ARA from sampling an empty layer range
The start index was sampled from [0, L//2] and the exclusive end index
from [L//2, L], so both could land on L//2 and leave no layers to
modify, wasting the trial. Starting the end range at L//2 + 1 rules
this out while keeping the bounds fixed, as multivariate TPE requires.
On the 2-layer test models, every ARA trial had picked this empty
range, so the seed-oss and gpt-oss tests never ran the optimization.
They now modify layer 1, which changes their expected hashes. The
seed-oss output also differs between the CPUs of GitHub's runners,
because L-BFGS amplifies small floating-point differences, so it gets
CI hashes as well.
* Re-initialize ARA's lora_A deterministically
ARA optimizes both LoRA matrices in place, but the model reset between
trials only zeroes lora_B. Each trial therefore started its
optimization from whatever lora_A the previous trial left behind.
Restoring a trial after the study re-optimized it from a different
starting point, so the exported model was not the one that earned the
trial's scores, and reproduction could not match hashes.
Re-initialize lora_A for every optimized module from a generator
seeded with the run seed. This makes the starting point independent of
trial history, the device and the global RNG state.
* Zero lora_A when resetting the model
The fast reset path only zeroed lora_B, and a fresh adapter gets lora_A
from PEFT's random init. Modules a trial doesn't modify therefore kept
whatever lora_A an earlier trial left, or a random one. The model is
unaffected since B is zero, but saved adapters differed byte-for-byte
depending on trial history, breaking hash verification for adapter
exports.
Zero lora_A as well, both on reset and right after applying LoRA.
* feat: add modifier base class
* feat: re-implement abliteration as a modifier plugin
* fix: adjust good/bad prompts hack to match tests
* feat: support dataset specifications containing multiple individual datasets
* feat: re-implement ARA as a modifier plugin
Arbitrary-Rank Ablation (ARA) (Weidmann 2026) was originally introduced by @p-e-w in #211
Co-authored-by: kabachuha <[email protected]>
Co-authored-by: joninco <[email protected]>
Co-authored-by: Ashar <[email protected]>
* feat: add tests for ARA
* feat: reduce default number of trials
* ci: remove broken "semantic-pull-request" workflow
* ci: add yet another alternative model hash
---------
Co-authored-by: kabachuha <[email protected]>
Co-authored-by: joninco <[email protected]>
Co-authored-by: Ashar <[email protected]>
* fix: Check response prefix.
In case if the model adds <think> at the end of user prompt and only generates </think> end, then the detection goes wrong.
* fix: Handle whitespace in CoT, remove redundant checks
* fix: Type checker
* fix: try looking at tag positions, not text end
regex
* ruff, fix import sorting
* fix: rich markup
* fix: I missed other ones
* feat: Update SHA256SUMS file hashes in the tests.
A major change that affects reproducibility.
* fix: It is now sensible to also update the extra SHA256SUMS.ci2 file.
* fix: Only consider whitespace, no other text or instructions.
because mistral-3 as additional reasoning instructions in its chat template. And I suppose many other models can have it too.
* fix: Update windows hashes.
* fix: Update CI hashes too
* as always update case two of mistral-3 (ci2 hash)
* docs: update comment
* feat: Handle the edge case for models having additional instructions.
* fix: Update windows hash for mistral-3
* docs: remove a line from the comments
because I'm not sure about GPT-OSS models' thinking tags and it cannot be confirmed using an untrained tiny GPT-OSS model. And inference fallback would of course generate gibberish as the model cannot understand additional instructions about 'how to generate response and how to think' from the chat_template.
* fix: a few things.
* fix: Update hash for qwen3.5 after the whitespace fix for its response prefix.
* docs: Update comment
* fix: Update qwen3.5 hash for CI
* fix: Remove Case 2 which only serves tests
unnecessary
* fix: Hash
* fix: concern is valid enough, so we use a small text.
add a comment too
* docs: minor
* feat: Allow specifying a specific config/subset name for the datasets.
This would be useful for using a single dataset that has harmful/harmless prompt pairs in different languages stored in different configs/subsets.
* fix: setting config/subset value when loading the dataset.
* fix: minor changes
* feat: add `reproducible` property to plugins
* feat: hard stop reproducibility and actually check `reproducible` field in the evaluator
* chore: flip default
* fix: check if plugin is built-in in repro gate
* fix: ruff
* fix: kld should be reproducible
* feat: move reproducible field to scorer level
* chore: move reproducible field back to plugin level
---------
Co-authored-by: mad-cat-lon <[email protected]>
* fix: validate scorer instance names in config
* fix: run scorer config tests in CI
* fix: drop redundant UV_PYTHON env and quote unittest pattern
* fix: report precise scorer name validation errors
* fix: minor change
use `W_org` matrix where needed...
* Update model.py
* Update model.py
* fix: Windows hash, remove BOM marker
* docs: Add info about test cases
* feat: Tests for row_normalization PRE & NONE
* feat: CI hash files for row_normalization PRE & NONE models
* feat: Documentation instructions about test suite
* add recommendation
* fix: remove notebook input shims
Closes#280
* feat: support headless operation (no interactive input)
* fix: prevent infinite loops
* feat: add end-to-end tests
* ci: run tests in CI
* ci: fix test output ordering
* fix: replace home-cooked `set_seed` function with Transformers builtin
* feat: print PyTorch config when running tests
* feat: print additional information
* experiment: try to standardize test environment
* fix: revert environment changes
* feat: support multiple valid hashes for each output file
* feat: add test output hashes for CI
* feat: add test output hashes for CI (alternative environment)
* feat: add hashes for Windows (#394)
* fix: Hash on windows
* trigger ci
* fix: prefer .yaml (used widely than .toml for model configs)
* use removeprefix
* docs: restore commet
* use removeprefix again
* tests: Add windows hash files for all test models
* trigger ci
* fix: minor cleanup
* clean merge mismatch
* remove unnecessary CRLF replace, now that we support more SUMS files
* fix: use binary mode for hashes everywhere
---------
Co-authored-by: Vinay Umrethe <[email protected]>
* fix: ensure utf-8 encoding for standard output and error to prevent UnicodeEncodeError on Windows
* fix: address bot review feedback
* refactor: deduplicate stream reconfiguration loop
* feat: let the optimizer disable MLP ablation via a 0 max_weight floor
The MLP max_weight lower bound was 0.8 for every component, so the optimizer
always applied at least 0.8x MLP ablation and could never turn it off, even
when ablating the MLP is pure collateral damage. Give the MLP a 0 lower bound
so the optimizer can disable it per model; attention keeps the 0.8 floor.
See #202.
* perf: skip the abliteration decomposition when the weight is 0
With a 0 max_weight the component's ablation is a no-op, and reset_model()
has already left the adapter at identity. Abort that layer/component before
the decomposition, which avoids the wasted work (and the degenerate
zero-matrix decomposition raised in review on #387).
* fix: clamp a negative MLP max_weight floor so 0 is reachable
A continuous suggest_float never samples exactly 0, so a 0 lower bound could
not actually disable the MLP. Use a small negative lower bound and clamp with
max(0, ...), which puts finite probability mass on exactly 0.
When a study is cancelled mid-way and the user selects 'Run additional
trials', settings.n_trials was incremented by n_additional_trials,
accumulating the original total into the new count. E.g. cancelling 200
trials at 30 and adding 10 gave n_trials=210 instead of 40, causing
'Running trial 31 of 210...' and planning 180 more trials instead of 10.
Fix by recalculating n_trials from actual completed trials + additional,
so the total reflects the new intended target, not the old one.
Fixes#379
Co-authored-by: Claude <[email protected]>
* feat: load reproduction information
* feat: check reproduction environment against original environment
* fix: remove `trust_remote_code` setting
This improves security when running Heretic with an untrusted config file. The prompt is now always shown.
This is NOT a breaking change, because we currently ignore values for unknown settings, so existing configs continue to work.
* feat: reproduce model from JSON file
* feat: verify hashes of uploaded weight files
* fix: fix issues in automatic reproduction system (#352)
* fix: Check if a model is gated / accessible
* fix: handle unknown gated models
* feat: Auto install requirements
* simplify
* Revert "simplify"
This reverts commit 10287926e9.
* Revert "feat: Auto install requirements"
This reverts commit f4be1abd04.
* fix: Seed pytorch method
* reference, style
* simplify token
* feat: Export strategy in reproduce.json, v2
* style: Name
* simplify export strategy
* style: Rename
* enumeration
* maybe remove seed as well
* fix: don't lock settings with permanent strategy
* simplify no choice, use try/finally block
* feat: verify hashes of locally saved weight files
* fix: remove obsolete code from merge
* docs: add automatic reproduction instructions to reproduce README
---------
Co-authored-by: Vinay-Umrethe <[email protected]>
* fix: make reset_model null-safe to handle study cancellations (#77)
* fix: address bot review, use nested getattr and fallback to settings dtypes
* fix: address maintainer review comments in model.py
* fix: address maintainer review feedback on reset_model
* fix: update Model.dtype type annotation to torch.dtype
* chore: revert pyproject.toml and uv.lock changes
* fix: fall back to exception class name when string representation is empty (#146)
* fix: walk stacktrace and causal chain to extract exception details in format_exception
* fix: fall back to complete stacktrace when exception has no message, as suggested by maintainer
* fix: address maintainer review, push newline control to printing boundaries
* feat: save processor for multimodal models
VL models load via AutoModelForImageTextToText, but only the tokenizer was
saved/pushed, dropping the processor's image/audio preprocessing config.
Save/push it alongside the tokenizer so multimodal models stay complete.
* Update src/heretic/model.py
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
* Adjusted processor type to use ProcessorMixin
---------
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
hf-transfer is declared in pyproject.toml but never activated: nothing in
the codebase sets HF_HUB_ENABLE_HF_TRANSFER, and downloads go through
from_pretrained / hf_hub_download with no transfer toggle. huggingface-hub
is pinned ~=1.7, where Xet is the default transfer backend, so hf-transfer
is dead weight and only surfaces a deprecation warning.
A dataset path that points to a plain file is now read as one prompt per
line, with empty lines ignored. For text files, "column" is ignored and
"split" is optional; when given, it selects a subset of lines using slice
notation (e.g. "[:400]").
Detection uses os.path.isfile so files without an extension also work. The
split-parsing logic is factored into a shared get_split_slice helper, which
derives the split name from the specification, and split/column are now
optional in DatasetSpecification, with the dataset branches raising a clear
error when either is missing. An invalid split raises instead of being
silently ignored.
A bare slice does not parse with the pinned datasets version, since
ReadInstruction.from_spec expects a named split, so the text branch prepends
a synthetic split name.
Revives the approach from #103.
Closes#98.
Co-authored-by: Ric <[email protected]>