Skip to content

[Common] Custom Annotation For Weights Compression - #4171

Draft
anzr299 wants to merge 46 commits into
openvinotoolkit:developfrom
anzr299:custom_annotation
Draft

anzr299 wants to merge 46 commits into
openvinotoolkit:developfrom
anzr299:custom_annotation

Conversation

@anzr299

@anzr299 anzr299 commented Aug 13, 2026 •

Copy link
Copy Markdown
Collaborator

Changes

New public API:

  • nncf.CustomAnnotation(scope, config): assigns a WeightCompressionConfig to a CustomAnnotationScope.
  • nncf.CustomAnnotationScope(names/patterns/types/subgraphs/validate): same as IgnoredScope. Both now derive from a shared BaseScope parent in scopes.py.
  • compress_weights(..., custom_annotation=[nncf.CustomAnnotation(scope1, config1), nncf.CustomAnnotation(scope2, config2)]).

Annotations are applied in WeightCompression.apply_custom_annotation() after the
configurations are assigned and before the weight parameters are fed to the main algorithms. modifying the parameters at this stage modifies it for the entirety of compression.

The precedence rules are:

  • An annotation overrides mixed precision, all_layers, backup_mode and ignored_scope.
  • If several annotations match the same layer, the last one in the list wins.
  • Configurations are keyed by weight name, so annotating any of them applies to the weight.
  • An annotation whose mode is not compressible is dropped, and the algorithm's
    configuration is kept.
    In all the above cases, appropriate warnings are given

Reason for changes

Use more precise compression schemes.

Related tickets

180366

Tests

Three tests were added to template_test_weights_compression.py

  • test_custom_annotation: the produced model configs are checked against reference configs.
  • test_custom_annotation_warning: the warnings are checked.
  • test_invalid_custom_annotation: the validation errors are checked.

The existing AWQ and scale estimation tests were extended to run on a custom-annotated
model as well.

Test examples - success

Weight compression - success

anzr299 added 3 commits August 5, 2026 12:14
…4164)

### Changes

Add TorchAO to GPTQModel Test Requirement and pin it to `0.17.0`

### Reason for changes

It was installed transitively before with torch. Recent release of
torchao(`0.18.0`) was being installed this way which caused unexpected
error. (Failing action:
https://github.com/openvinotoolkit/nncf/actions/runs/30928990933/job/92221632392)

### Related tickets

<!--- Post the numerical ID of the ticket, if available -->

### Tests

<!--- How was the correctness of changes tested and whether new tests
were added -->

[Test
examples](https://github.com/openvinotoolkit/nncf/actions/runs/30984266172)
- success
@github-actions github-actions Bot added documentation Improvements or additions to documentation NNCF OpenVINO Pull requests that updates NNCF OpenVINO API Public API-impacting changes labels Aug 13, 2026
@github-actions github-actions Bot added the NNCF Common Pull request that updates NNCF Common label Aug 13, 2026
@github-actions github-actions Bot removed the NNCF OpenVINO Pull requests that updates NNCF OpenVINO label Sep 15, 2026
@github-actions github-actions Bot added NNCF PT Pull requests that updates NNCF PyTorch NNCF OpenVINO Pull requests that updates NNCF OpenVINO NNCF ONNX Pull requests that updates NNCF ONNX labels Sep 15, 2026
@anzr299
anzr299 marked this pull request as draft September 23, 2026 20:24

This comment was marked as resolved.

This comment was marked as resolved.

This comment was marked as resolved.

This comment was marked as resolved.

This comment was marked as resolved.

mode=CompressWeightsMode.INT4_SYM,
group_size=64,
ratio=1.0,
custom_annotation=[

@AlexanderDokuchaev AlexanderDokuchaev Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In my mind, add 3 new classes it's too mach for API
WeightCompressionConfig and BaseScope are just dataclasses, wich can be comvined in one

from nncf.scopes import BaseScope
from nncf.quantization.algorithms.weight_compression.config import WeightCompressionConfig

class ScopedConfig(WeightCompressionConfig, BaseScope):
    pass

def compress_weights(
   ...
    scoped_configs: list[ScopedConfig] | None = None,
) -> TModel:

It can be used as it in existed function, or can create instance of scope and config by adding methos to_scope() and to_config()

import nncf
from nncf import compress_weights, ScopedConfig

compressed_model = compress_weights(
    model, # model is openvino.Model object
    mode=CompressWeightsMode.INT4_SYM,
    group_size=64,
    ratio=1.0,
    scoped_configs=[  # or  `override_configs` 
        ScopedConfig(patterns=[".*self_attn.*", ".*router.*"]. mode=CompressWeightsMode.INT8_ASYM, group_size=-1),
        ScopedConfig(patterns=[".*mul.*"]. mode=CompressWeightsMode.INT4_ASYM, group_size=128)
    ],
)

model, # model is openvino.Model object
mode=CompressWeightsMode.INT4_SYM,
group_size=64,
ratio=1.0,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How ratio will works? How many layers will be compressed to backaup mode, if setup custom annotation?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ratio will work like it works if we for example use ignored scope.
The custom annotated layers are not considered for ratio defining params. Ratio = 0.8 means 80% of ratio defining parameters but excluding the custom annotated ndoes.


- The scope of an annotation is defined by the same rules as [the ignored scope](/docs/usage/IgnoredScope.md): node names,
regular expressions, operation types and subgraphs.
The configuration given by an annotation takes precedence over the decision made by the algorithm, namely over the

@AlexanderDokuchaev AlexanderDokuchaev Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thinks that we should simply behavoir in this case.
Now it's hard to describe and undestand how it should works: custom_config, ignoredscope, mixed precision, backup mode, ration ..

My suggestion
Ordering of selection of config.

  1. Ignored scope - find layers that will be stay as it
  2. custom annotation - select config by manual selection (the first item in list has higher priority)
  3. all layers that no mattached with ignored and custom annotation scopes got config from main api (mode, group_size, ..)
  4. all other parameters (ratio, backup_mode, ..) ignores with warning or better with exception

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

API Public API-impacting changes documentation Improvements or additions to documentation NNCF Common Pull request that updates NNCF Common NNCF ONNX Pull requests that updates NNCF ONNX NNCF OpenVINO Pull requests that updates NNCF OpenVINO NNCF PT Pull requests that updates NNCF PyTorch

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants