0% found this document useful (0 votes)

18 views32 pages

Webir 06

This lecture focuses on the Vector Space Model for web information retrieval, discussing how documents can be represented as vectors using tf-idf values. It emphasizes the importance of cosine similarity over Euclidean distance for measuring document proximity to queries, which are also treated as vectors. The lecture concludes with considerations for integrating vector space models with various query types and scoring methods.

Uploaded by

Tilahun Eshetu

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

18 views32 pages

Webir 06

Uploaded by

Tilahun Eshetu

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

You are on page 1/ 32

Web Information Retrieval

Lecture 6
Vector Space Model
Recap of the last lecture
 Parametric and field searches
 Zones in documents
 Scoring documents: zone weighting
 Index support for scoring
 tfidf and vector spaces
This lecture
 Vector space model
 Efficiency considerations
 Nearest neighbors and approximations
Documents as vectors
 At the end of Lecture 5 we said:
 Each doc j can now be viewed as a vector of tfidf
values, one component for each term
 So we have a vector space
 terms are axes
 docs live in this space
 even with stemming, may have 20,000+ dimensions
Example

Antony and Cleopatra Julius Caesar The Tempest Hamlet Othello Macbeth

Brutus 3.0 8.3 0.0 1.0 0.0 0.0

Caesar 2.3 2.3 0.0 0.5 0.3 0.3
mercy 0.5 0.0 0.7 0.9 0.9 0.3
Why turn docs into vectors?
 First application: Query-by-example
 Given a doc D, find others “like” it.
 Now that D is a vector, find vectors (docs) “near” it.
Intuition
t3
d2

d3
d1

d5
t2
d4

Postulate: Documents that are “close together”

in the vector space talk about the same things.
The vector space model
Query as vector:
 We regard query as short document

 We return the documents ranked by the closeness of

their vectors to the query, also represented as a

vector.
Desiderata for proximity
 If d1 is near d2, then d2 is near d1.
 If d1 near d2, and d2 near d3, then d1 is not far from d3.
 No doc is closer to d than d itself.
First cut
 Distance between d1 and d2 is the length of the vector
|d1 – d2|.
 Euclidean distance
 Why is this not a great idea?
 We still haven’t dealt with the issue of length
normalization
 However, we can implicitly normalize by looking at
angles instead
Sec. 6.3

Why distance is a bad idea

The Euclidean
distance between q
and d2 is large even
though the
distribution of terms
in the query q and
the distribution of
terms in the
document d2 are very
similar.
Sec. 6.3

Use angle instead of distance

 Thought experiment: take a document d and append it to
itself. Call this document d′.
 “Semantically” d and d′ have the same content
 The Euclidean distance between the two documents can
be quite large
 The angle between the two documents is 0,
corresponding to maximal similarity.

 Key idea: Rank documents according to angle with

query.
Sec. 6.3

From angles to cosines

 The following two notions are equivalent.
 Rank documents in decreasing order of the angle between
query and document
 Rank documents in increasing order of
cosine(query,document)
 Cosine is a monotonically decreasing function for the
interval of interest [0o, 90o]
Sec. 6.3

From angles to cosines

 But how – and why – should we be computing cosines?

Cosine similarity
 Distance between vectors d1 and d2 captured by the
cosine of the angle x between them.
 Note – this is similarity, not distance

t3
d2

d1
θ

t2
Cosine similarity
 A vector can be normalized (given a length of 1) by
dividing each of its components by its length – here
we use the L2 norm
x2 x x
i
2
i

 This maps vectors onto the unit sphere:



M
 Then, dj  i 1
wi , j  1
 Longer documents don’t get more weight
Cosine similarity
 

M
d j  dk wi , j wi ,k
sim(d j , d k )  cos(d j , d k )     i 1

i 1 w i1 i,k
M M
d j dk 2
i, j w 2

 Cosine of angle between two vectors

 The denominator involves the lengths of the vectors.

Normalization
Normalized vectors
 For normalized vectors, the cosine is simply the dot
product:
   
cos(d j , d k )  d j  d k
Cosine similarity exercises
 Exercise: Rank the following by decreasing cosine
similarity:
 Two docs that have only frequent words (the, a, an, of)
in common.
 Two docs that have no words in common.
 Two docs that have many rare words in common
(wingspan, tailfin).
Exercise
 Euclidean distance between vectors:

 d  d i ,k 
M
d j  dk 
2
i 1 i, j

 Show that, for normalized vectors, Euclidean

distance gives the same proximity ordering as the
cosine measure
Example
 Docs: Austen's Sense and Sensibility, Pride and
Prejudice; Bronte's Wuthering Heights
SaS PaP WH
affection 115 58 20
jealous 10 7 11
gossip 2 0 6

SaS PaP WH
affection 0.996 0.993 0.847
jealous 0.087 0.120 0.466
gossip 0.017 0.000 0.254
Example
 Docs: Austen's Sense and Sensibility, Pride and
Prejudice; Bronte's Wuthering Heights
SaS PaP WH
affection 115 58 20
jealous 10 7 11
gossip 2 0 6

SaS PaP WH
affection 0.996 0.993 0.847
jealous 0.087 0.120 0.466
gossip 0.017 0.000 0.254

 cos(SAS, PAP) = .996 x .993 + .087 x .120 + .017 x 0.0 = 0.999

 cos(SAS, WH) = .996 x .847 + .087 x .466 + .017 x .254 = 0.889
Queries as vectors
 Key idea 1: Do the same for queries: represent them
as vectors in the space
 Key idea 2: Rank documents according to their
proximity to the query in this space
 proximity = similarity of vectors
Cosine(query,document)
Unit vectors
Dot product
  


M
  qd q d qi d i
cos(q , d )         i 1
q d
i 1 q i1 i
M M
qd 2
i d 2

cos(q, d) is the cosine similarity of q and d or, equivalently,

the cosine of the angle between q and d.

Summary: What’s the real point of
using vector spaces?
 Key: A user’s query can be viewed as a (very) short
document.
 Query becomes a vector in the same space as the
docs.
 Can measure each doc’s proximity to it.
 Natural measure of scores/ranking – no longer
Boolean.
 Queries are expressed as bags of words
 Other similarity measures: see
http://www.lans.ece.utexas.edu/~strehl/diss/node52.html for a
survey
Interaction: vectors and phrases
 Phrases don’t fit naturally into the vector space world:
 “hong kong” “new york”
 Positional indexes don’t capture tf/idf information for
“hong kong”
 Biword indexes treat certain phrases as terms
 For these, can pre-compute tf/idf.
 A hack: we cannot expect end-user formulating
queries to know what phrases are indexed
Vectors and Boolean queries
 Vectors and Boolean queries really don’t work
together very well
 We cannot express AND, OR, NOT, just by summing
term frequencies
Vector spaces and other operators
 Vector space queries are apt for no-syntax, bag-of-
words queries
 Clean metaphor for similar-document queries
 Not a good combination with Boolean, positional
query operators, phrase queries, …
 But …
Query language vs. scoring
 May allow user a certain query language, say
 Freetext basic queries
 Phrase, wildcard etc. in Advanced Queries.
 For scoring (oblivious to user) may use all of the
above, e.g. for a freetext query
 Highest-ranked hits have query as a phrase
 Next, docs that have all query terms near each other
 Then, docs that have some query terms, or all of them
spread out, with tf x idf weights for scoring
Exercises
 How would you augment the inverted index built in
lectures 1–3 to support cosine ranking computations?
 What information do we need to store?
 Walk through the steps of serving a query.
 The math of the vector space model is quite
straightforward, but being able to do cosine ranking
efficiently at runtime is nontrivial
Resources
 IIR Chapters 6.3, 7.3

L14 VSM
No ratings yet
L14 VSM
24 pages
TF-IDF and Ranked Retrieval Basics
No ratings yet
TF-IDF and Ranked Retrieval Basics
51 pages
ISR Chap... 5
No ratings yet
ISR Chap... 5
34 pages
Vector Space Model: TF - IDF: Adapted From Lectures by
No ratings yet
Vector Space Model: TF - IDF: Adapted From Lectures by
37 pages
Chapter 4 - Part II
No ratings yet
Chapter 4 - Part II
44 pages
Lec 3
No ratings yet
Lec 3
51 pages
Vector Space Model
No ratings yet
Vector Space Model
11 pages
IR Lecture 4b
No ratings yet
IR Lecture 4b
57 pages
L04
No ratings yet
L04
35 pages
Boolean and Vector Space Retrieval Models
No ratings yet
Boolean and Vector Space Retrieval Models
33 pages
TF Idf
100% (3)
TF Idf
38 pages
IR Lecture 4b
No ratings yet
IR Lecture 4b
57 pages
Module 3 Indexing Part A
No ratings yet
Module 3 Indexing Part A
46 pages
Boolean and Vector Space Retrieval Models
No ratings yet
Boolean and Vector Space Retrieval Models
27 pages
06 VectorSpaceModel
No ratings yet
06 VectorSpaceModel
65 pages
Vector Space Model
No ratings yet
Vector Space Model
7 pages
06 VectorSpaceModel PDF
No ratings yet
06 VectorSpaceModel PDF
75 pages
Session 4 Text Feature
No ratings yet
Session 4 Text Feature
40 pages
Lecture 04
No ratings yet
Lecture 04
41 pages
AI6122 Topic 3.2 - Ranking
No ratings yet
AI6122 Topic 3.2 - Ranking
27 pages
Lec2 2
No ratings yet
Lec2 2
17 pages
Module-7 Similarity Measure
No ratings yet
Module-7 Similarity Measure
39 pages
Boolean and Vector Space Retrieval Models
No ratings yet
Boolean and Vector Space Retrieval Models
31 pages
Chapter 6 - Scoring Term Weighting and Vector Space Model
No ratings yet
Chapter 6 - Scoring Term Weighting and Vector Space Model
43 pages
Text
No ratings yet
Text
11 pages
Vector Space Model
No ratings yet
Vector Space Model
4 pages
Lecture 3 VSM
No ratings yet
Lecture 3 VSM
16 pages
Vector Space Model for IR Students
No ratings yet
Vector Space Model for IR Students
23 pages
Precision Recal TF Idf
No ratings yet
Precision Recal TF Idf
36 pages
Term Weighting & The Vector Space Model
No ratings yet
Term Weighting & The Vector Space Model
2 pages
CS583 Info Retrieval
No ratings yet
CS583 Info Retrieval
34 pages
Tif Idf Cosine Similarity
No ratings yet
Tif Idf Cosine Similarity
44 pages
Frontiers of Computational Journalism - Columbia Journalism School Fall 2012 - Week 3: Document Topic Modeling
No ratings yet
Frontiers of Computational Journalism - Columbia Journalism School Fall 2012 - Week 3: Document Topic Modeling
48 pages
Similarity Measures Le 512
No ratings yet
Similarity Measures Le 512
14 pages
Lecture 5 - Scoring, Term Weighting, Vector Space Model - Part 1
No ratings yet
Lecture 5 - Scoring, Term Weighting, Vector Space Model - Part 1
45 pages
Lecture 10
No ratings yet
Lecture 10
18 pages
Retrieval Models & Ranking Overview
No ratings yet
Retrieval Models & Ranking Overview
16 pages
Web and Traditional IR Methods
No ratings yet
Web and Traditional IR Methods
24 pages
Vector Space Model & Tf-idf Explained
100% (1)
Vector Space Model & Tf-idf Explained
16 pages
Chapter 4 IR Models
No ratings yet
Chapter 4 IR Models
43 pages
Modern Information Retrieval Chapter 5 Query Operations
No ratings yet
Modern Information Retrieval Chapter 5 Query Operations
33 pages
Relevance of A Document To A Query
No ratings yet
Relevance of A Document To A Query
10 pages
Chapter 5 IR
No ratings yet
Chapter 5 IR
46 pages
L02-IR Models MMN
No ratings yet
L02-IR Models MMN
27 pages
Lecture 05
No ratings yet
Lecture 05
51 pages
Module 1 Part BInformation Retrieval Webdocuments
No ratings yet
Module 1 Part BInformation Retrieval Webdocuments
49 pages
Ranked Retrieval
No ratings yet
Ranked Retrieval
52 pages
Data Mining: Similarity and Distance Recommendation Systems Sketching, Locality Sensitive Hashing
No ratings yet
Data Mining: Similarity and Distance Recommendation Systems Sketching, Locality Sensitive Hashing
57 pages
Chapter Five IR Models
No ratings yet
Chapter Five IR Models
28 pages
Machine Learning For Natural Language Processing: Classification: Nearest Neighbors
No ratings yet
Machine Learning For Natural Language Processing: Classification: Nearest Neighbors
28 pages
3 Retrieval Models
No ratings yet
3 Retrieval Models
87 pages
Vector Space Model
No ratings yet
Vector Space Model
10 pages
5 IRModels IR
No ratings yet
5 IRModels IR
25 pages
IR Models for Students
No ratings yet
IR Models for Students
62 pages
CS583 Info Retrieval
No ratings yet
CS583 Info Retrieval
33 pages
Vector Space Model
No ratings yet
Vector Space Model
11 pages
Chapter Five Conflict and Conflict Management
No ratings yet
Chapter Five Conflict and Conflict Management
17 pages
Medical Laboratory
No ratings yet
Medical Laboratory
17 pages
1.para Introduction
No ratings yet
1.para Introduction
52 pages
CH 2
No ratings yet
CH 2
70 pages
Information Retrieval Systems
No ratings yet
Information Retrieval Systems
46 pages
Statment Debreberhan
No ratings yet
Statment Debreberhan
12 pages
OM For MBA1
No ratings yet
OM For MBA1
297 pages
Presentation1 Training
No ratings yet
Presentation1 Training
27 pages
Exam 1-1-1-1
No ratings yet
Exam 1-1-1-1
4 pages
Exam 3-1-1-1
No ratings yet
Exam 3-1-1-1
4 pages
Leve 2 Coc Exam
No ratings yet
Leve 2 Coc Exam
43 pages
Exam 4-1
No ratings yet
Exam 4-1
5 pages
MBA OR CH3 Final C
No ratings yet
MBA OR CH3 Final C
105 pages
Accounting Coc Level 1
No ratings yet
Accounting Coc Level 1
8 pages
Mba or CH 1 Final C
No ratings yet
Mba or CH 1 Final C
62 pages
4model Exam Foundation Engineering II
100% (1)
4model Exam Foundation Engineering II
5 pages
MBA OR CH21 Final C
No ratings yet
MBA OR CH21 Final C
82 pages
Northwest Fire District: Balanced Scorecard Strategic Planning & Management System
No ratings yet
Northwest Fire District: Balanced Scorecard Strategic Planning & Management System
38 pages
Worksheet For Business Research Methods
No ratings yet
Worksheet For Business Research Methods
2 pages
Work Sheet For Business Law
No ratings yet
Work Sheet For Business Law
3 pages
Module For Managerial Economics
No ratings yet
Module For Managerial Economics
136 pages
Work Sheet For System Analysis and Design
No ratings yet
Work Sheet For System Analysis and Design
5 pages
Strategy Map Evaluation Checklist
No ratings yet
Strategy Map Evaluation Checklist
1 page
Motivation & Communication
No ratings yet
Motivation & Communication
75 pages
Informal Business Reports Guide
No ratings yet
Informal Business Reports Guide
5 pages
Introduction To Business Communication
No ratings yet
Introduction To Business Communication
43 pages
Planning Function & Decision-Making
No ratings yet
Planning Function & Decision-Making
38 pages
A Preliminary Concept Meaning & Importance of Controlling The Controlling Process Types of Control Control Techniques/methods
No ratings yet
A Preliminary Concept Meaning & Importance of Controlling The Controlling Process Types of Control Control Techniques/methods
20 pages
Mba CH4
No ratings yet
Mba CH4
65 pages
Mba CH1
No ratings yet
Mba CH1
32 pages
Module Reading Writing Quarter 4
No ratings yet
Module Reading Writing Quarter 4
94 pages
Grade 12 Research Module
No ratings yet
Grade 12 Research Module
20 pages
Final Exam Opeartion Research
No ratings yet
Final Exam Opeartion Research
1 page
Tajamul Resume
No ratings yet
Tajamul Resume
3 pages
Cree LED JSeries Feature Sheet
No ratings yet
Cree LED JSeries Feature Sheet
2 pages
Monster Hunter 2 Damage
No ratings yet
Monster Hunter 2 Damage
48 pages
Pharmacist Return To Work Course
100% (2)
Pharmacist Return To Work Course
8 pages
CH 11. Mensuration 1
No ratings yet
CH 11. Mensuration 1
21 pages
Digitalization Impact on Bank Service Efficiency
No ratings yet
Digitalization Impact on Bank Service Efficiency
7 pages
LMI P1 Series Parts List Chemical Metering Pump PDF
No ratings yet
LMI P1 Series Parts List Chemical Metering Pump PDF
4 pages
Research Ii Espine, de Los Santos, Verzo Effects of Technology To Senior Hi
No ratings yet
Research Ii Espine, de Los Santos, Verzo Effects of Technology To Senior Hi
46 pages
Procurement Summary Report
No ratings yet
Procurement Summary Report
3 pages
Agura, Danieli - Bullet Journal8
No ratings yet
Agura, Danieli - Bullet Journal8
2 pages
Philosophical Foundation
No ratings yet
Philosophical Foundation
2 pages
5 - Chloroplast Development in Green Plant Tissues The Interplay Between Light Hormone
No ratings yet
5 - Chloroplast Development in Green Plant Tissues The Interplay Between Light Hormone
17 pages
2nd Sem 1st Periodical Exam CPAR 2022-2023
100% (1)
2nd Sem 1st Periodical Exam CPAR 2022-2023
4 pages
Moblie Attachment
No ratings yet
Moblie Attachment
52 pages
20224WEMT CourseGrade5Week11Thinking SkillsCommon
No ratings yet
20224WEMT CourseGrade5Week11Thinking SkillsCommon
5 pages
How To Use TP5100 2A 8.4 - 4.2V 1S and 2S Lithium Battery Charger
No ratings yet
How To Use TP5100 2A 8.4 - 4.2V 1S and 2S Lithium Battery Charger
5 pages
Raymundo Vs Luneta Motor 58 Phil 889
No ratings yet
Raymundo Vs Luneta Motor 58 Phil 889
4 pages
Accounting for Capital Changes
No ratings yet
Accounting for Capital Changes
7 pages
The Magic Garden of George B and Other Logic Puzzles 1st Edition Raymond Smullyan PDF Download
100% (10)
The Magic Garden of George B and Other Logic Puzzles 1st Edition Raymond Smullyan PDF Download
61 pages
Probability: Experience
No ratings yet
Probability: Experience
12 pages
English Grammar Error Detection & Correction
No ratings yet
English Grammar Error Detection & Correction
231 pages
FSP 150CC-GE110 Series: Compact Carrier Ethernet Service Demarcation at The Edge
No ratings yet
FSP 150CC-GE110 Series: Compact Carrier Ethernet Service Demarcation at The Edge
4 pages
Advanced Linear Algebra Guide
100% (1)
Advanced Linear Algebra Guide
270 pages
Automated Stock Management System
No ratings yet
Automated Stock Management System
7 pages
Lesson 1 - PE10 - Fitness
100% (1)
Lesson 1 - PE10 - Fitness
37 pages
22 November 2024
No ratings yet
22 November 2024
15 pages
Chap 002
No ratings yet
Chap 002
16 pages

Webir 06

Uploaded by

Webir 06

Uploaded by

Web Information Retrieval

Brutus 3.0 8.3 0.0 1.0 0.0 0.0

Postulate: Documents that are “close together”

 We return the documents ranked by the closeness of

their vectors to the query, also represented as a

Why distance is a bad idea

Use angle instead of distance

 Key idea: Rank documents according to angle with

From angles to cosines

From angles to cosines

 But how – and why – should we be computing cosines?

 This maps vectors onto the unit sphere:

 Cosine of angle between two vectors

 Show that, for normalized vectors, Euclidean

 cos(SAS, PAP) = .996 x .993 + .087 x .120 + .017 x 0.0 = 0.999

cos(q, d) is the cosine similarity of q and d or, equivalently,

the cosine of the angle between q and d.

You might also like