Conversation
…4164) ### Changes Add TorchAO to GPTQModel Test Requirement and pin it to `0.17.0` ### Reason for changes It was installed transitively before with torch. Recent release of torchao(`0.18.0`) was being installed this way which caused unexpected error. (Failing action: https://github.com/openvinotoolkit/nncf/actions/runs/30928990933/job/92221632392) ### Related tickets <!--- Post the numerical ID of the ticket, if available --> ### Tests <!--- How was the correctness of changes tested and whether new tests were added --> [Test examples](https://github.com/openvinotoolkit/nncf/actions/runs/30984266172) - success
…o an/fx/conformance/fix
| mode=CompressWeightsMode.INT4_SYM, | ||
| group_size=64, | ||
| ratio=1.0, | ||
| custom_annotation=[ |
There was a problem hiding this comment.
In my mind, add 3 new classes it's too mach for API
WeightCompressionConfig and BaseScope are just dataclasses, wich can be comvined in one
from nncf.scopes import BaseScope
from nncf.quantization.algorithms.weight_compression.config import WeightCompressionConfig
class ScopedConfig(WeightCompressionConfig, BaseScope):
pass
def compress_weights(
...
scoped_configs: list[ScopedConfig] | None = None,
) -> TModel:It can be used as it in existed function, or can create instance of scope and config by adding methos to_scope() and to_config()
import nncf
from nncf import compress_weights, ScopedConfig
compressed_model = compress_weights(
model, # model is openvino.Model object
mode=CompressWeightsMode.INT4_SYM,
group_size=64,
ratio=1.0,
scoped_configs=[ # or `override_configs`
ScopedConfig(patterns=[".*self_attn.*", ".*router.*"]. mode=CompressWeightsMode.INT8_ASYM, group_size=-1),
ScopedConfig(patterns=[".*mul.*"]. mode=CompressWeightsMode.INT4_ASYM, group_size=128)
],
)| model, # model is openvino.Model object | ||
| mode=CompressWeightsMode.INT4_SYM, | ||
| group_size=64, | ||
| ratio=1.0, |
There was a problem hiding this comment.
How ratio will works? How many layers will be compressed to backaup mode, if setup custom annotation?
There was a problem hiding this comment.
ratio will work like it works if we for example use ignored scope.
The custom annotated layers are not considered for ratio defining params. Ratio = 0.8 means 80% of ratio defining parameters but excluding the custom annotated ndoes.
|
|
||
| - The scope of an annotation is defined by the same rules as [the ignored scope](/docs/usage/IgnoredScope.md): node names, | ||
| regular expressions, operation types and subgraphs. | ||
| The configuration given by an annotation takes precedence over the decision made by the algorithm, namely over the |
There was a problem hiding this comment.
I thinks that we should simply behavoir in this case.
Now it's hard to describe and undestand how it should works: custom_config, ignoredscope, mixed precision, backup mode, ration ..
My suggestion
Ordering of selection of config.
- Ignored scope - find layers that will be stay as it
- custom annotation - select config by manual selection (the first item in list has higher priority)
- all layers that no mattached with ignored and custom annotation scopes got config from main api (mode, group_size, ..)
- all other parameters (ratio, backup_mode, ..) ignores with warning or better with exception
Changes
New public API:
nncf.CustomAnnotation(scope, config): assigns aWeightCompressionConfigto a CustomAnnotationScope.nncf.CustomAnnotationScope(names/patterns/types/subgraphs/validate): same asIgnoredScope. Both now derive from a sharedBaseScopeparent inscopes.py.compress_weights(..., custom_annotation=[nncf.CustomAnnotation(scope1, config1), nncf.CustomAnnotation(scope2, config2)]).Annotations are applied in
WeightCompression.apply_custom_annotation()after theconfigurations are assigned and before the weight parameters are fed to the main algorithms. modifying the parameters at this stage modifies it for the entirety of compression.
The precedence rules are:
all_layers,backup_modeandignored_scope.configuration is kept.
In all the above cases, appropriate warnings are given
Reason for changes
Use more precise compression schemes.
Related tickets
180366
Tests
Three tests were added to
template_test_weights_compression.pytest_custom_annotation: the produced model configs are checked against reference configs.test_custom_annotation_warning: the warnings are checked.test_invalid_custom_annotation: the validation errors are checked.The existing AWQ and scale estimation tests were extended to run on a custom-annotated
model as well.
Test examples - success
Weight compression - success