Semester project 3 for Linguistics
- analyze some text
- learn techniques for analyzing text
- learn some software engineering skills
- learn git, markdown, use LaTeX
-
who chooses the task? you or me?
-
how many tasks? 1, 2 or 4
-
how many teams? 1, 2 or 4
-
Introduce yourselves
- name
- languages spoken (natural/programming)
- interests (relevant to the course)
- one fun fact (not relevant to the course)
- your choice!
- LLM text vs Human text
- what is different (test for Czech?)
- I have data for English
- Distribution of lexicalized metaphor
- combine sense-tagged text with ChainNet
- analyze distribution
- fill in what is missing
- Improve Czech wordnet
- fill in Czech specific patterns
- adverbs [of languages]
- diminutives
- aspect variants
- gender variants (role nouns)
- generate definitions
- get examples from corpora
- derivational links with affixes?
- fill in Czech specific patterns
- get something useful for each task
- combine to make a best-of-breed
- write, submit and publish a paper
- release at least one automatically tagged, aligned corpus
- use github to coordinate
-
15-30 minutes progress
-
longer discussion of issues as necessary
-
small meetings pair+me or WSD+me, .... as necessary
- ALL make github account
- send me accountname
- I will add to github