Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 6 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ Some of sensAI's key benefits are:
* **A unifying interface to a wide variety of model classes across frameworks**

Apply the same principles to a wide variety of models, whether they are
neural networks, tree ensembles or non-parametric models – without
neural networks, tree ensembles or non-parametric models — without
losing the ability of exploiting each model's particular strengths.

sensAI supports models based on PyTorch, scikit-learn, XGBoost and
Expand All @@ -36,7 +36,7 @@ Some of sensAI's key benefits are:

Modularise data pre-processing steps and features generation, representing
the properties of features explicitly.
* For each model, select a suitable subset of features, composing the
* For each model, select a suitable subset of features, composing
the desired feature generators in order to obtain an initial
input pipeline.
* Transform the features into representations that are optimised for
Expand Down Expand Up @@ -220,7 +220,7 @@ features_df = feature_collector.get_multi_feature_generator().generate(df)

Depending on the type of model, the representation of the input data may need to
be adapted. For instance,
some models can directly process arbitarily represented categorical data, others
some models can directly process arbitrarily represented categorical data, others
require an encoding. Some models can deal with arbitrary scales of numerical
data, others work best with normalised data.

Expand Down Expand Up @@ -350,7 +350,7 @@ tensor-based representations). See our tutorial on neural network models.
### Evaluation

Evaluating the performance of models can be a chore.
sensAI's high-level evaluation classes severely cut down on the boiler plate,
sensAI's high-level evaluation classes severely cut down on the boilerplate,
allowing you to focus on what matters.

```
Expand Down Expand Up @@ -438,7 +438,7 @@ be overlooked.

sensAI supports combinatorial optimisation via

* **stochastic local search**, provding implementations of
* **stochastic local search**, providing implementations of
* simulated annealing
* parallel tempering.

Expand All @@ -462,7 +462,7 @@ sensAI's `util` package contains a wide range of general utilities, including

# Documentation

* [Reference documentation and tutorials](https://opcode81.github.io/sensAI/docs/)
* [Reference documentation and tutorials](https://oraios.github.io/sensAI/docs/)

At this point, the documentation is still limited, but we plan to add
further tutorials and overview documentation in the future.
Expand Down
4 changes: 2 additions & 2 deletions docs/1-supervised-learning/1-vector-models.ipynb
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@
"## VectorModel\n",
"\n",
"The backbone of supervised learning implementations is the `VectorModel` abstraction.\n",
"It is so named, because, in computer science, a *vector* corresponds to an array of data,\n",
"It is so named because in computer science a *vector* corresponds to an array of data,\n",
"and vector models map such vectors to the desired outputs, i.e. regression targets or \n",
"classes.\n",
"\n",
Expand All @@ -76,7 +76,7 @@
"\n",
"### DataFrame-Based Interfaces\n",
"\n",
"Vector models use pandas DataFrames as the fundmental input and output data structures.\n",
"Vector models use pandas DataFrames as the fundamental input and output data structures.\n",
"Every row in a data frame corresponds to a vector of data, and an entire data frame can thus be viewed as a dataset or batch of data. Data frames are a good base representation for input data because\n",
" * they provide rudimentary meta-data in the form of column names, avoiding ambiguity.\n",
" * they can contain arbitrarily complex data, yet in the simplest of cases, they can directly be mapped to a data matrix (2D array) of features that simple models can directly process.\n",
Expand Down
2 changes: 1 addition & 1 deletion docs/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ Some of sensAI's key benefits are:
* **A unifying interface to a wide variety of model classes across frameworks**

Apply the same principles to a wide variety of models, whether they are
neural networks, tree ensembles or non-parametric models – without
neural networks, tree ensembles or non-parametric models — without
losing the ability of exploiting each model's particular strengths.

sensAI supports models based on PyTorch, scikit-learn, XGBoost and
Expand Down
12 changes: 8 additions & 4 deletions notebooks/intro.ipynb
Original file line number Diff line number Diff line change
Expand Up @@ -16,14 +16,18 @@
"%autoreload 2"
],
"metadata": {
"tags": ["hide-cell"]
"tags": [
"hide-cell"
]
}
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"tags": ["hide-cell"]
"tags": [
"hide-cell"
]
},
"outputs": [],
"source": [
Expand Down Expand Up @@ -553,15 +557,15 @@
" \n",
" Whatever the case may be, we can represent it in a data frame. We call the original input data frame, which we pass to a sensAI ``VectorModel``, the *raw data frame*.\n",
"\n",
" 2. **Extracing features from the raw data, using their \"natural\" representation** (using ``FeatureGenerators``)\n",
" 2. **Extracting features from the raw data, using their \"natural\" representation** (using ``FeatureGenerators``)\n",
" \n",
" We extract from the raw data frame pieces of information that we regard as relevant *features* for the task at hand.\n",
" A sensAI ``FeatureGenerator`` can generate one or more data frame columns (containing arbitrary data), and a model can be associated with any number of feature generators.\n",
" Several key aspects:\n",
"\n",
" * FeatureGenerators crititcally decouple the original raw data from the features used by the models, enabling different models to use different sets of features or \n",
" entirely different representations of the same features.\n",
" * FeatureGenerators become part of the model and are (where necessary) jointly trained with model. This facilitates model deployment, as every sensAI model becomes a single unit\n",
" * FeatureGenerators become part of the model and are (where necessary) jointly trained with the model. This facilitates model deployment, as every sensAI model becomes a single unit\n",
" that can directly process raw input data, which is (usually) straightforward to supply at inference time.\n",
" * FeatureGenerators store meta-data on the features they generate, enabling downstream components to handle them appropriately.\n",
"\n",
Expand Down
10 changes: 6 additions & 4 deletions notebooks/neural_networks.ipynb
Original file line number Diff line number Diff line change
Expand Up @@ -33,9 +33,11 @@
{
"cell_type": "code",
"execution_count": null,
"metadata": {"tags": [
"hide-cell"
]},
"metadata": {
"tags": [
"hide-cell"
]
},
"outputs": [],
"source": [
"%%capture\n",
Expand Down Expand Up @@ -177,7 +179,7 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"Neural networks work best on **normalised inputs**, so we have opted to apply basic normalisation by specifying a normalisation mode which will transforms inputs by dividing by the maximum value found across all columns in the training data. For more elaborate normalisation options, we could have used a data frame transformer (DFT), particularly `DFTNormalisation` or `DFTSkLearnTransformer`.\n",
"Neural networks work best on **normalised inputs**, so we have opted to apply basic normalisation by specifying a normalisation mode which will transform inputs by dividing by the maximum value found across all columns in the training data. For more elaborate normalisation options, we could have used a data frame transformer (DFT), particularly `DFTNormalisation` or `DFTSkLearnTransformer`.\n",
"\n",
"sensAI's default **neural network training algorithm** is based on early stopping, which involves checking, in regular intervals, the performance of the model on a validation set (which is split from the training set) and ultimately selecting the model that performed best on the validation set. You have full control over the loss evaluation method used to select the best model (by passing a respective `NNLossEvaluator` instance to NNOptimiserParams) as well as the method that is used to split the training set into the actual training set and the validation set (by adding a `DataFrameSplitter` to the model or using a custom `TorchDataSetProvider`).\n",
"\n",
Expand Down
2 changes: 1 addition & 1 deletion src/sensai/tracking/tracking_base.py
Original file line number Diff line number Diff line change
Expand Up @@ -198,7 +198,7 @@ def begin_optional_tracking_context_for_model(self, model: VectorModelBase, trac
Furthermore, tracking can be disabled by passing `track=False` even if a tracked experiment is present.

:param model: the model for which to begin tracking
:paraqm track: whether tracking shall be enabled; if False, force use of a dummy context which performs no actual tracking even
:param track: whether tracking shall be enabled; if False, force use of a dummy context which performs no actual tracking even
if a tracked experiment is present
:return: a context manager that can be used to track results for the given model
"""
Expand Down