Skip to content

[FEATURE REQUEST] Functionality for analyzing the differences between two Annotation objects. #67

Description

@AndriyMulyar

What problem does your feature solve?
A method to do analysis of annotations (namely for the application of looking at differences between gold and predicted annotations).

Describe the solution you'd like
The Annotation class should be given some static methods like Annotation.diff(ann_object_1, ann_object_2) will output the difference between to annotation objects. Maybe some parameter for leniency to deal with fuzzy annotation matching.

Interface sklearn to compute various evaluation metrics between two annotation files (assuming one is gold and one is predicted).

Additional context
This would be very useful for result analysis and guiding the building of pipelines.

Activity

  1. AndriyMulyar commented on Jan 1, 2019

    @AndriyMulyar
    CollaboratorAuthor

    Currently pull request #68 begins preliminary work mentioned above.

    Ideas for further improvements:

    1. Method in Dataset that will allow to compare gold and predicted over a whole corpus by utilizing the diff functionality implemented in Added Annotations.diff(); unit tests #68 .
    2. Give the diff method optional fuzzy parameters that will highlight model predictions that are almost correct (maybe off by a few characters)
  2. swfarnsworth commented on Jan 10, 2019

    @swfarnsworth
    Member

    I believe the functionality you described is covered by Annotations.compare_by_index(), which has the strict parameter for fuzzy predictions.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions