Repository navigation
[Common] Custom Annotation For Weights Compression #4171
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: develop
Are you sure you want to change the base?
Changes from all commits
7dd77a2
ada24e3
0d930a5
feb3cfa
5a63d44
99aad99
7e75fd2
8fddfd1
feb8417
bd4cea8
bb9ffbd
1dd2c06
dc94fe3
fffc8cf
b2da0b4
cb96a10
dc1f780
f2c9905
c9c483b
47fdcd5
a9545cc
1f3b327
feb1563
85719b5
75e8225
d510074
c090461
e87a62e
dd6587c
cd62da1
a177728
481cc68
1f93893
2bc0e20
0f35160
1018999
78bc8b7
0366367
ee2f437
a231b6a
73a39dc
bc8c923
d77e634
83690da
e93e574
48b3f6c
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -108,6 +108,43 @@ from nncf import compress_weights, CompressWeightsMode | |
| compressed_model = compress_weights(model, mode=CompressWeightsMode.INT4_ASYM, group_size=64, ratio=0.9) # model is openvino.Model object | ||
| ``` | ||
|
|
||
| #### Custom precision for specific layers | ||
|
|
||
| - The `mode`, `ratio` and `backup_mode` parameters define the precision of a layer indirectly: the mixed-precision | ||
| algorithm decides which layers are compressed to the primary precision, and the rest is compressed to the backup one. | ||
| The `custom_annotation` parameter makes it possible to assign a compression configuration to certain layers | ||
| explicitly. A typical use case is a Mixture-of-Experts model, where the experts are compressed to 4 bits, while the | ||
| attention and the router layers are kept in 8 bits to preserve accuracy. | ||
|
|
||
| ```python | ||
| import nncf | ||
| from nncf import compress_weights, CompressWeightsMode | ||
|
|
||
| compressed_model = compress_weights( | ||
| model, # model is openvino.Model object | ||
| mode=CompressWeightsMode.INT4_SYM, | ||
| group_size=64, | ||
| ratio=1.0, | ||
| custom_annotation=[ | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. In my mind, add 3 new classes it's too mach for API from nncf.scopes import BaseScope
from nncf.quantization.algorithms.weight_compression.config import WeightCompressionConfig
class ScopedConfig(WeightCompressionConfig, BaseScope):
pass
def compress_weights(
...
scoped_configs: list[ScopedConfig] | None = None,
) -> TModel:It can be used as it in existed function, or can create instance of scope and config by adding methos import nncf
from nncf import compress_weights, ScopedConfig
compressed_model = compress_weights(
model, # model is openvino.Model object
mode=CompressWeightsMode.INT4_SYM,
group_size=64,
ratio=1.0,
scoped_configs=[ # or `override_configs`
ScopedConfig(patterns=[".*self_attn.*", ".*router.*"]. mode=CompressWeightsMode.INT8_ASYM, group_size=-1),
ScopedConfig(patterns=[".*mul.*"]. mode=CompressWeightsMode.INT4_ASYM, group_size=128)
],
) |
||
| nncf.CustomAnnotation( | ||
| scope=nncf.CustomAnnotationScope(patterns=[".*self_attn.*", ".*router.*"]), | ||
| config=nncf.WeightCompressionConfig(mode=CompressWeightsMode.INT8_ASYM, group_size=-1), | ||
| ), | ||
| ], | ||
| ) | ||
| ``` | ||
|
|
||
| - The scope of an annotation is defined by the same rules as [the ignored scope](/docs/usage/IgnoredScope.md): node names, | ||
| regular expressions, operation types and subgraphs. | ||
| The configuration given by an annotation takes precedence over the decision made by the algorithm, namely over the | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I thinks that we should simply behavoir in this case. My suggestion
|
||
| mixed-precision assignment, the `all_layers` option and the backup precision of the embeddings and the last linear | ||
| layer. A layer that is matched by both the `ignored_scope` and a custom annotation is compressed with the | ||
| user-defined configuration and a warning is logged. If several annotations match the same layer, the last one takes | ||
| precedence. A weight shared by several layers is compressed once, so annotating any of them applies the | ||
| configuration to this weight. The codebook compression modes are not supported by an annotation. The mode of an | ||
| annotation must be supported by the backend, the same as the `mode` option, e.g. NF4 can not be annotated for a | ||
| Torch, TorchFX or ONNX model. | ||
|
|
||
| #### Data-aware methods | ||
|
|
||
| - Accuracy of the 4-bit compressed models can be improved by using data-aware mixed-precision algorithm. It is capable to find outliers in the input activations and assign different quantization precision to minimize accuracy degradation. | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
How ratio will works? How many layers will be compressed to backaup mode, if setup custom annotation?
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
ratio will work like it works if we for example use ignored scope.
The custom annotated layers are not considered for ratio defining params. Ratio = 0.8 means 80% of ratio defining parameters but excluding the custom annotated ndoes.