Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
46 commits
Select commit Hold shift + click to select a range
7dd77a2
[Example] Add TorchAO to GPTQModel Test Requirement (#4164)
anzr299 Aug 5, 2026
ada24e3
Merge branch 'develop' of https://github.com/openvinotoolkit/nncf int…
anzr299 Aug 13, 2026
0d930a5
init
anzr299 Aug 13, 2026
feb3cfa
change ingored scope util function names to common name
anzr299 Aug 13, 2026
5a63d44
group size fallback takes precedence over custom annotation
anzr299 Aug 13, 2026
99aad99
Merge branch 'openvinotoolkit:develop' into custom_annotation
anzr299 Aug 27, 2026
7e75fd2
fix int8 group size
anzr299 Aug 31, 2026
8fddfd1
Merge branch 'develop' into custom_annotation
anzr299 Sep 10, 2026
feb8417
fix failing precommit test
anzr299 Sep 10, 2026
bd4cea8
Merge branch 'develop' into custom_annotation
anzr299 Sep 15, 2026
bb9ffbd
update tests
anzr299 Sep 15, 2026
1dd2c06
upd
anzr299 Sep 15, 2026
dc94fe3
update refs
anzr299 Sep 15, 2026
fffc8cf
newline
anzr299 Sep 15, 2026
b2da0b4
ref
anzr299 Sep 15, 2026
cb96a10
test
anzr299 Sep 15, 2026
dc1f780
use full ref
anzr299 Sep 15, 2026
f2c9905
refactor
anzr299 Sep 15, 2026
c9c483b
refactor ordering
anzr299 Sep 16, 2026
47fdcd5
ref
anzr299 Sep 16, 2026
a9545cc
remove old scope kind
anzr299 Sep 18, 2026
1f3b327
add shared weight edge case
anzr299 Sep 18, 2026
feb1563
Merge branch 'develop' into custom_annotation
anzr299 Sep 22, 2026
85719b5
fix
anzr299 Sep 22, 2026
75e8225
clean
anzr299 Sep 22, 2026
d510074
fix
anzr299 Sep 22, 2026
c090461
shared weights also give precendence to custom annotation
anzr299 Sep 22, 2026
e87a62e
fix multiple warnings; third shared node case
anzr299 Sep 23, 2026
dd6587c
extend custom annotation to awq and scale estimation tests
anzr299 Sep 23, 2026
cd62da1
fix docstring
anzr299 Sep 23, 2026
a177728
rework
anzr299 Sep 23, 2026
481cc68
fix refs
anzr299 Sep 23, 2026
1f93893
Merge branch 'develop' into custom_annotation
anzr299 Sep 23, 2026
2bc0e20
fix docstring
anzr299 Sep 23, 2026
0f35160
remove support for codebook; support edge case with shared weight whe…
anzr299 Sep 23, 2026
1018999
update comments
anzr299 Sep 23, 2026
78bc8b7
backend specific mode validity check
anzr299 Sep 23, 2026
0366367
support backup mode none with data aware compression
anzr299 Sep 24, 2026
ee2f437
remove remaining codebook support
anzr299 Sep 24, 2026
a231b6a
validate group sizes correctly for mx, nv datatypes
anzr299 Sep 24, 2026
73a39dc
get the group size appropriately at WC COnfig intialization
anzr299 Sep 24, 2026
bc8c923
add error for group size fallback with MX, NV types for common function
anzr299 Sep 24, 2026
d77e634
add validation for the openvino backend case.
anzr299 Sep 24, 2026
83690da
remove warning for nodes without weight.
anzr299 Sep 24, 2026
e93e574
validate custom annotation with advanced parameters
anzr299 Sep 24, 2026
48b3f6c
revert path + [Key]
anzr299 Sep 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -108,6 +108,43 @@ from nncf import compress_weights, CompressWeightsMode
compressed_model = compress_weights(model, mode=CompressWeightsMode.INT4_ASYM, group_size=64, ratio=0.9) # model is openvino.Model object
```

#### Custom precision for specific layers

- The `mode`, `ratio` and `backup_mode` parameters define the precision of a layer indirectly: the mixed-precision
algorithm decides which layers are compressed to the primary precision, and the rest is compressed to the backup one.
The `custom_annotation` parameter makes it possible to assign a compression configuration to certain layers
explicitly. A typical use case is a Mixture-of-Experts model, where the experts are compressed to 4 bits, while the
attention and the router layers are kept in 8 bits to preserve accuracy.

```python
import nncf
from nncf import compress_weights, CompressWeightsMode

compressed_model = compress_weights(
model, # model is openvino.Model object
mode=CompressWeightsMode.INT4_SYM,
group_size=64,
ratio=1.0,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How ratio will works? How many layers will be compressed to backaup mode, if setup custom annotation?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ratio will work like it works if we for example use ignored scope.
The custom annotated layers are not considered for ratio defining params. Ratio = 0.8 means 80% of ratio defining parameters but excluding the custom annotated ndoes.

custom_annotation=[

@AlexanderDokuchaev AlexanderDokuchaev Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In my mind, add 3 new classes it's too mach for API
WeightCompressionConfig and BaseScope are just dataclasses, wich can be comvined in one

from nncf.scopes import BaseScope
from nncf.quantization.algorithms.weight_compression.config import WeightCompressionConfig

class ScopedConfig(WeightCompressionConfig, BaseScope):
    pass

def compress_weights(
   ...
    scoped_configs: list[ScopedConfig] | None = None,
) -> TModel:

It can be used as it in existed function, or can create instance of scope and config by adding methos to_scope() and to_config()

import nncf
from nncf import compress_weights, ScopedConfig

compressed_model = compress_weights(
    model, # model is openvino.Model object
    mode=CompressWeightsMode.INT4_SYM,
    group_size=64,
    ratio=1.0,
    scoped_configs=[  # or  `override_configs` 
        ScopedConfig(patterns=[".*self_attn.*", ".*router.*"]. mode=CompressWeightsMode.INT8_ASYM, group_size=-1),
        ScopedConfig(patterns=[".*mul.*"]. mode=CompressWeightsMode.INT4_ASYM, group_size=128)
    ],
)

nncf.CustomAnnotation(
scope=nncf.CustomAnnotationScope(patterns=[".*self_attn.*", ".*router.*"]),
config=nncf.WeightCompressionConfig(mode=CompressWeightsMode.INT8_ASYM, group_size=-1),
),
],
)
```

- The scope of an annotation is defined by the same rules as [the ignored scope](/docs/usage/IgnoredScope.md): node names,
regular expressions, operation types and subgraphs.
The configuration given by an annotation takes precedence over the decision made by the algorithm, namely over the

@AlexanderDokuchaev AlexanderDokuchaev Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thinks that we should simply behavoir in this case.
Now it's hard to describe and undestand how it should works: custom_config, ignoredscope, mixed precision, backup mode, ration ..

My suggestion
Ordering of selection of config.

  1. Ignored scope - find layers that will be stay as it
  2. custom annotation - select config by manual selection (the first item in list has higher priority)
  3. all layers that no mattached with ignored and custom annotation scopes got config from main api (mode, group_size, ..)
  4. all other parameters (ratio, backup_mode, ..) ignores with warning or better with exception

mixed-precision assignment, the `all_layers` option and the backup precision of the embeddings and the last linear
layer. A layer that is matched by both the `ignored_scope` and a custom annotation is compressed with the
user-defined configuration and a warning is logged. If several annotations match the same layer, the last one takes
precedence. A weight shared by several layers is compressed once, so annotating any of them applies the
configuration to this weight. The codebook compression modes are not supported by an annotation. The mode of an
annotation must be supported by the backend, the same as the `mode` option, e.g. NF4 can not be annotated for a
Torch, TorchFX or ONNX model.

#### Data-aware methods

- Accuracy of the 4-bit compressed models can be improved by using data-aware mixed-precision algorithm. It is capable to find outliers in the input activations and assign different quantization precision to minimize accuracy degradation.
Expand Down
3 changes: 3 additions & 0 deletions src/nncf/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -65,8 +65,11 @@
from nncf.quantization.advanced_parameters import AdvancedQuantizationParameters as AdvancedQuantizationParameters
from nncf.quantization.advanced_parameters import AdvancedScaleEstimationParameters as AdvancedScaleEstimationParameters
from nncf.quantization.advanced_parameters import AdvancedSmoothQuantParameters as AdvancedSmoothQuantParameters
from nncf.quantization.advanced_parameters import CustomAnnotation as CustomAnnotation
from nncf.quantization.advanced_parameters import GroupSizeFallbackMode as GroupSizeFallbackMode
from nncf.quantization.advanced_parameters import OverflowFix as OverflowFix
from nncf.quantization.algorithms.weight_compression.config import WeightCompressionConfig as WeightCompressionConfig
from nncf.scopes import CustomAnnotationScope as CustomAnnotationScope
from nncf.scopes import IgnoredScope as IgnoredScope
from nncf.scopes import Subgraph as Subgraph
from nncf.version import __version__ as __version__
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,7 @@
from nncf.parameters import TargetDevice
from nncf.quantization.advanced_parameters import AdvancedCompressionParameters
from nncf.quantization.advanced_parameters import AdvancedQuantizationParameters
from nncf.quantization.advanced_parameters import CustomAnnotation
from nncf.quantization.algorithms.post_training.algorithm import PostTrainingQuantization
from nncf.quantization.algorithms.weight_compression.algorithm import WeightCompression
from nncf.scopes import IgnoredScope
Expand Down Expand Up @@ -132,6 +133,7 @@ def compress_weights_impl(
backup_mode: BackupMode,
compression_format: CompressionFormat,
advanced_parameters: AdvancedCompressionParameters | None = None,
custom_annotation: list[CustomAnnotation] | None = None,
) -> torch.fx.GraphModule:
"""
Implementation of the `compress_weights()` method for the Torch Fx backend.
Expand All @@ -151,6 +153,7 @@ def compress_weights_impl(
backup_mode,
compression_format,
advanced_parameters,
custom_annotation,
)
graph = build_graph(model)
compressed_model = compression_algorithm.apply(model, graph, dataset=dataset)
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@
from nncf.experimental.torch.sparsify_activations.target_scope import TargetScope
from nncf.experimental.torch.sparsify_activations.target_scope import get_target_node_names_from_target_scope
from nncf.scopes import IgnoredScope
from nncf.scopes import get_ignored_node_names_from_ignored_scope
from nncf.scopes import get_node_names_from_scope
from nncf.torch.model_creation import is_wrapped_model
from nncf.torch.model_creation import wrap_model

Expand Down Expand Up @@ -181,9 +181,7 @@ def _get_target_sparsity_by_node(self, graph: NNCFGraph) -> dict[NNCFNode, float
:return: A dictionary with nodes and the corresponding target sparsity level.
"""
supported_metatypes = self._backend_entity.supported_metatypes
ignored_names = get_ignored_node_names_from_ignored_scope(
self._ignored_scope, graph, strict=self._ignored_scope.validate
)
ignored_names = get_node_names_from_scope(self._ignored_scope, graph, strict=self._ignored_scope.validate)
target_sparsity_by_node = {}
for scope, target_sparsity in self._target_sparsity_by_scope.items():
target_names = get_target_node_names_from_target_scope(scope, graph, strict=scope.validate)
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -14,8 +14,8 @@
import nncf
from nncf.common.graph.graph import NNCFGraph
from nncf.scopes import IgnoredScope
from nncf.scopes import get_difference_ignored_scope
from nncf.scopes import get_matched_ignored_scope_info
from nncf.scopes import get_difference_scope
from nncf.scopes import get_matched_scope_info


@dataclass
Expand Down Expand Up @@ -84,7 +84,7 @@ def get_target_node_names_from_target_scope(
:param strict: Whether target_scope must match at least one node or not.
:return: NNCF node names from the given graph matched by target scope.
"""
matched_target_scope, matches = get_matched_ignored_scope_info(target_scope, [nncf_graph])
matched_target_scope, matches = get_matched_scope_info(target_scope, [nncf_graph])
if strict:
_check_target_scope_strictly_matched(target_scope, matched_target_scope)
return set().union(*matches.values())
Expand All @@ -97,7 +97,7 @@ def _check_target_scope_strictly_matched(target_scope: TargetScope, matched_targ
:param target_scope: The given target scope.
:param matched_target_scope: The actual target scope matched in a graph.
"""
unmatched_scope = get_difference_ignored_scope(target_scope, matched_target_scope)
unmatched_scope = get_difference_scope(target_scope, matched_target_scope)
error_messages = []
for match_type in ("names", "types", "patterns", "subgraphs"):
unmatched_rules = getattr(unmatched_scope, match_type)
Expand Down
3 changes: 3 additions & 0 deletions src/nncf/onnx/quantization/quantize_model.py
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@
from nncf.quantization.advanced_parameters import AdvancedAccuracyRestorerParameters
from nncf.quantization.advanced_parameters import AdvancedCompressionParameters
from nncf.quantization.advanced_parameters import AdvancedQuantizationParameters
from nncf.quantization.advanced_parameters import CustomAnnotation
from nncf.quantization.advanced_parameters import QuantizationParameters
from nncf.quantization.algorithms.accuracy_control.algorithm import QuantizationAccuracyRestorer
from nncf.quantization.algorithms.accuracy_control.algorithm import calculate_accuracy_drop
Expand Down Expand Up @@ -329,6 +330,7 @@ def compress_weights_impl(
backup_mode: BackupMode,
compression_format: CompressionFormat,
advanced_parameters: AdvancedCompressionParameters | None = None,
custom_annotation: list[CustomAnnotation] | None = None,
) -> onnx.ModelProto:
if model.opset_import[0].version < 13:
msg = "ONNX models with opset version < 13 do not support per-channel quantization."
Expand Down Expand Up @@ -361,6 +363,7 @@ def compress_weights_impl(
backup_mode,
compression_format,
advanced_parameters,
custom_annotation,
)
graph = build_graph(model)

Expand Down
7 changes: 5 additions & 2 deletions src/nncf/openvino/quantization/quantize_model.py
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,7 @@
from nncf.quantization.advanced_parameters import AdvancedAccuracyRestorerParameters
from nncf.quantization.advanced_parameters import AdvancedCompressionParameters
from nncf.quantization.advanced_parameters import AdvancedQuantizationParameters
from nncf.quantization.advanced_parameters import CustomAnnotation
from nncf.quantization.advanced_parameters import convert_to_dict_recursively
from nncf.quantization.algorithms.accuracy_control.algorithm import QuantizationAccuracyRestorer
from nncf.quantization.algorithms.accuracy_control.algorithm import calculate_accuracy_drop
Expand All @@ -55,7 +56,7 @@
from nncf.quantization.statistics_caching import cache_weight_compression_statistics
from nncf.quantization.statistics_caching import register_statistics_for_algorithm
from nncf.scopes import IgnoredScope
from nncf.scopes import validate_ignored_scope
from nncf.scopes import validate_scope

TTensor = TypeVar("TTensor")

Expand Down Expand Up @@ -96,7 +97,7 @@ def _extract_all_subgraphs(model: ov.Model, current_id: str) -> None:
main_model_graph_id = "main_model_graph"
_extract_all_subgraphs(model, main_model_graph_id)
if ignored_scope and ignored_scope.validate:
validate_ignored_scope(ignored_scope, graphs.values())
validate_scope(ignored_scope, graphs.values())
ignored_scope = IgnoredScope(
ignored_scope.names, ignored_scope.patterns, ignored_scope.types, ignored_scope.subgraphs, validate=False
)
Expand Down Expand Up @@ -379,6 +380,7 @@ def compress_weights_impl(
backup_mode: BackupMode,
compression_format: CompressionFormat,
advanced_parameters: AdvancedCompressionParameters | None = None,
custom_annotation: list[CustomAnnotation] | None = None,
) -> ov.Model:
"""
Implementation of the `compress_weights()` method for the OpenVINO backend.
Expand All @@ -400,6 +402,7 @@ def compress_weights_impl(
backup_mode,
compression_format,
advanced_parameters,
custom_annotation,
Comment thread
anzr299 marked this conversation as resolved.
)

statistics_points = None
Expand Down
28 changes: 26 additions & 2 deletions src/nncf/openvino/rt_info.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,9 @@

import nncf
from nncf.common.logging import nncf_logger
from nncf.scopes import IgnoredScope
from nncf.quantization.advanced_parameters import CustomAnnotation
from nncf.quantization.algorithms.weight_compression.config import WeightCompressionConfig
from nncf.scopes import BaseScope


def exclude_empty_fields(value: dict[str, Any]) -> dict[str, Any]:
Expand All @@ -33,6 +35,16 @@ def exclude_empty_fields(value: dict[str, Any]) -> dict[str, Any]:
return value


def compression_config_to_dict(config: WeightCompressionConfig) -> dict[str, Any]:
"""
Converts a weight compression config into a dictionary suitable for dumping into Model's meta section.

:param config: Weight compression config.
:return: Dictionary with the compression config parameters.
"""
return {"mode": config.mode.value, "group_size": config.group_size}


def dump_parameters(
model: ov.Model, parameters: dict[str, Any], algo_name: str | None = "quantization", path: list[str] | None = None
) -> None:
Expand All @@ -48,13 +60,25 @@ def dump_parameters(
path = path if path else []
for key, value in parameters.items():
# Special condition for composed fields like IgnoredScope
if isinstance(value, IgnoredScope):
if isinstance(value, BaseScope):
value = exclude_empty_fields(asdict(value))
if bool(value):
dump_parameters(model, value, algo_name, [key])
continue
# The default value in case empty ignored_scope parameter passed
value = []
# Special condition for the list of custom annotations
elif isinstance(value, (list, tuple)) and any(isinstance(item, CustomAnnotation) for item in value):
for i, annotation in enumerate(value):
# Index each annotation so its scope and config stay paired and annotations
# do not overwrite each other, e.g. custom_annotation/0/{scope/patterns,config}.
dump_parameters(
model,
{"scope": annotation.scope, "config": compression_config_to_dict(annotation.config)},
algo_name,
path + [key, str(i)],
)
continue

rt_path = ["nncf", algo_name] + path + [key]
model.set_rt_info(str(value), rt_path)
Expand Down
31 changes: 31 additions & 0 deletions src/nncf/quantization/advanced_parameters.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,9 @@
from nncf.common.quantization.structs import QuantizationScheme as QuantizationMode
from nncf.common.utils.api_marker import api
from nncf.parameters import StrEnum
from nncf.quantization.algorithms.weight_compression.config import WeightCompressionConfig
from nncf.quantization.range_estimator import RangeEstimatorParameters
from nncf.scopes import CustomAnnotationScope
from nncf.tensor import TensorDataType

TTensor = Any
Expand Down Expand Up @@ -451,6 +453,35 @@ class AdvancedCompressionParameters:
)


@api(canonical_alias="nncf.CustomAnnotation")
@dataclass
class CustomAnnotation:
"""
Applies a user-defined weight compression configuration to a portion of a model.
It overrides the configuration assigned by the weight compression algorithm, including the mixed precision
algorithm and the `ignored_scope`, `all_layers` and `backup_mode` options.

Example:

.. code-block:: python

import nncf

annotation = nncf.CustomAnnotation(
scope=nncf.CustomAnnotationScope(patterns=['.*self_attn.*', '.*router.*']),
config=nncf.WeightCompressionConfig(mode=nncf.CompressWeightsMode.INT8_ASYM, group_size=-1),
)

:param scope: Defines the portion of a model to annotate.
:type scope: nncf.CustomAnnotationScope
:param config: Weight compression configuration to apply to the matched nodes.
:type config: nncf.WeightCompressionConfig
"""

scope: CustomAnnotationScope = field(default_factory=CustomAnnotationScope)
config: WeightCompressionConfig = field(default_factory=WeightCompressionConfig)


@api()
@dataclass
class AdvancedAccuracyRestorerParameters:
Expand Down
4 changes: 2 additions & 2 deletions src/nncf/quantization/algorithms/min_max/algorithm.py
Original file line number Diff line number Diff line change
Expand Up @@ -69,7 +69,7 @@
from nncf.quantization.range_estimator import RangeEstimatorParametersSet
from nncf.quantization.range_estimator import StatisticsType
from nncf.scopes import IgnoredScope
from nncf.scopes import get_ignored_node_names_from_ignored_scope
from nncf.scopes import get_node_names_from_scope

TModel = TypeVar("TModel")

Expand Down Expand Up @@ -613,7 +613,7 @@ def _get_ignored_names(
:param ignored_patterns: Ignored patterns.
:return: Ignored node names and ignore reason for quantization.
"""
user_ignored_names = get_ignored_node_names_from_ignored_scope(
user_ignored_names = get_node_names_from_scope(
self._ignored_scope, nncf_graph, strict=self._ignored_scope.validate
)
autogenerated_ignored_names = get_node_names_matching_graph_pattern(inference_nncf_graph, ignored_patterns)
Expand Down
Loading
Loading