JP2009282593A - Method, server and program for managing index data for retrieving content - Google Patents
Method, server and program for managing index data for retrieving content Download PDFInfo
- Publication number
- JP2009282593A JP2009282593A JP2008131603A JP2008131603A JP2009282593A JP 2009282593 A JP2009282593 A JP 2009282593A JP 2008131603 A JP2008131603 A JP 2008131603A JP 2008131603 A JP2008131603 A JP 2008131603A JP 2009282593 A JP2009282593 A JP 2009282593A
- Authority
- JP
- Japan
- Prior art keywords
- users
- index data
- keyword
- content
- server
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Images
Landscapes
- Information Transfer Between Computers (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
【課題】コンテンツの識別データの有用性に基づく索引データを管理する方法、サーバ、およびプログラムを提供すること。
【解決手段】検索サーバ10は、Webページを閲覧したユーザにより当該Webページに関連付けられた、ブックマークを受信し、受信したブックマークに含まれる、所定の種類の語句の集合を抽出し、抽出された語句の集合と、当該語句の集合が抽出されたブックマークに関連付けられたWebページと、の同一の組み合わせについて、当該関連付けが行われたユーザ数を、当該組み合わせと対応付けて記憶し、記憶されたユーザ数に基づいて、当該ユーザ数に対応する組み合わせについて、語句の集合に含まれる語句それぞれをキーワードとした索引データを生成する。
【選択図】図3A method, server, and program for managing index data based on usefulness of identification data of content are provided.
A search server receives a bookmark associated with a Web page by a user viewing the Web page, extracts a set of words of a predetermined type included in the received bookmark, and is extracted For the same combination of a set of phrases and the Web page associated with the bookmark from which the set of phrases is extracted, the number of users associated is stored in association with the combination and stored. Based on the number of users, for each combination corresponding to the number of users, index data is generated using each of the phrases included in the set of phrases as a keyword.
[Selection] Figure 3
Description
本発明は、索引データを管理する方法、サーバ、およびプログラムに関する。 The present invention relates to a method, a server, and a program for managing index data.
従来、インターネットにおいてWebページや画像ファイル等で提供されるコンテンツは、多種多量であるため、必要な情報を得ようとするユーザは、検索システムを利用して、キーワードによる検索を行ってきた。この検索システムは、検索の効率化を図るために、予めキーワードとコンテンツとを関連付けた索引データを生成しておくことが多い。 Conventionally, since there are a large amount of contents provided as Web pages, image files, and the like on the Internet, a user who wants to obtain necessary information has performed a search using a keyword using a search system. In many cases, this search system generates index data in which keywords and contents are associated in advance in order to improve the search efficiency.
このような索引データには、コンテンツ内から抽出された単語や、コンテンツに付与されたタグがキーワードとして用いられる。これにより、ユーザは、コンテンツの内容を示すキーワードを入力することにより検索することができる。 In such index data, words extracted from the content and tags attached to the content are used as keywords. Thereby, the user can search by inputting the keyword which shows the content content.
また、ユーザが有用と認めたコンテンツに対しては、ブックマークを登録しておき、このブックマークを選択することにより、該当のコンテンツに素早く辿り着くようにすることも行われている。このブックマークは、ユーザが利用する端末装置に固有のものには限られず、所定のサーバにて複数ユーザが登録したブックマークが管理される場合もある。 In addition, a bookmark is registered for content recognized as useful by the user, and the user can quickly reach the corresponding content by selecting the bookmark. This bookmark is not limited to a terminal device used by the user, and bookmarks registered by a plurality of users may be managed by a predetermined server.
ところで、このようなブックマークが多数登録された場合には、所望のブックマーク自体が発見され難いため、ブックマークを検索して検索結果を提示する管理方法も提案されている(例えば、特許文献1参照)。そこで、ブックマークから抽出される単語を、コンテンツ内の単語やタグ等と同様に、索引データのキーワードとすることも考えられる。
しかしながら、コンテンツ内の単語や、タグ、ブックマーク等の識別データをキーワードとして検索を行う場合、抽出されたキーワードには、コンテンツの内容と関連度の低いものが含まれる場合がある。特に、ユーザにより指定されるタグやブックマークに基づくキーワードは、他のユーザからは分かり難いことも多い。このような関連度の低いキーワードにより検索の精度が低下したり、分かり難いキーワードにより無用の索引データが生成されたりすることが多かった。 However, when a search is performed using identification data such as words in content, tags, bookmarks, and the like as keywords, the extracted keywords may include those having a low degree of association with the content contents. In particular, tags specified by users and keywords based on bookmarks are often difficult for other users to understand. In many cases, such a low-relevance keyword reduces the accuracy of search, and unnecessary index data is generated by a keyword that is difficult to understand.
したがって、個々のキーワードの有用性を考慮して索引データを生成することが重要である。そこで、本発明は、コンテンツの識別データの有用性に基づく索引データを管理する方法、サーバ、およびプログラムを提供することを目的とする。 Therefore, it is important to generate index data in consideration of the usefulness of individual keywords. Therefore, an object of the present invention is to provide a method, a server, and a program for managing index data based on the usefulness of content identification data.
本発明では、以下のような解決手段を提供する。 The present invention provides the following solutions.
(1) キーワードと当該キーワードにより検索されるコンテンツとを対応付けた索引データをサーバが管理する方法であって、
ユーザにより当該コンテンツに関連付けられた、識別データを受信する受信ステップと、
前記受信ステップにより受信した識別データに含まれる、所定の種類の語句の集合を抽出する抽出ステップと、
前記抽出ステップにより抽出された語句の集合と、当該語句の集合が抽出された識別データに関連付けられたコンテンツと、の同一の組み合わせについて、当該関連付けが行われたユーザ数を、当該組み合わせと対応付けて記憶する記憶ステップと、
前記記憶ステップにより記憶された前記ユーザ数に基づいて、当該ユーザ数に対応する前記組み合わせについて、前記語句の集合に含まれる語句それぞれを前記キーワードとした前記索引データを生成する生成ステップと、を含む方法。
(1) A method in which a server manages index data in which a keyword is associated with content searched for by the keyword,
A receiving step of receiving identification data associated with the content by the user;
An extraction step of extracting a set of words of a predetermined type included in the identification data received by the reception step;
For the same combination of the set of phrases extracted by the extraction step and the content associated with the identification data from which the set of phrases is extracted, the number of users associated is associated with the combination. Memory step for storing
Generating based on the number of users stored in the storing step, the index data for the combination corresponding to the number of users, the index data using each of the words included in the set of words as the keyword. Method.
このような構成によれば、当該方法を実行するサーバは、コンテンツ(例えば、Webページや画像ファイル等)を閲覧したユーザにより当該コンテンツに関連付けられた、識別データ(例えば、ブックマークやタグ等)を受信し、受信した識別データに含まれる、所定の種類(例えば、名詞等)の語句の集合を抽出し、抽出された語句の集合と、当該語句の集合が抽出された識別データに関連付けられたコンテンツと、の同一の組み合わせについて、当該関連付けが行われたユーザ数を、当該組み合わせと対応付けて記憶し、記憶されたユーザ数に基づいて、当該ユーザ数に対応する組み合わせについて、語句の集合に含まれる語句それぞれをキーワードとした索引データを生成する。 According to such a configuration, the server that executes the method receives identification data (for example, a bookmark or a tag) associated with the content by a user who has viewed the content (for example, a web page or an image file). A set of phrases of a predetermined type (for example, nouns) included in the received identification data is extracted, and the set of extracted phrases and the set of phrases are associated with the extracted identification data For the same combination with content, the number of associated users is stored in association with the combination, and based on the stored number of users, the combination corresponding to the number of users is stored in a set of phrases. Index data using each included phrase as a keyword is generated.
このことにより、当該サーバは、ブックマークやタグ等の識別データに含まれるキーワードの集合を単位として、この集合が登録された数(ユーザ数)を記憶する。このユーザ数が多いことは、同一のキーワード集合が抽出される識別データを登録したユーザが多いことを示すため、このキーワード集合がコンテンツと深く関連し、有用なキーワードの組み合わせであると判断できる。そして、当該サーバは、この記憶されたユーザ数に基づいて、集合に含まれるキーワードに関する索引データを生成するので、多くのユーザに認知された有用なキーワードに基づいて索引データを生成することができる。 As a result, the server stores the number of registered groups (number of users) in units of keyword groups included in identification data such as bookmarks and tags. A large number of users indicates that there are many users who have registered identification data from which the same keyword set is extracted. Therefore, it can be determined that this keyword set is deeply related to the content and is a useful combination of keywords. And since the said server produces | generates the index data regarding the keyword contained in a set based on this memorize | stored number of users, it can produce | generate index data based on the useful keyword recognized by many users. .
したがって、この索引データを用いる検索システムは、コンテンツとの関連度の高い有用な索引により精度の高い検索を実行できる。したがって、ユーザは、要求に合った有用なコンテンツを効率的に抽出して閲覧できる可能性がある。 Therefore, the search system using the index data can execute a highly accurate search using a useful index having a high degree of association with the content. Therefore, the user may be able to efficiently extract and browse useful content that meets the requirements.
(2) 前記記憶ステップは、所定の期間に前記受信ステップにより受信した識別データについて、前記ユーザ数を記憶することを特徴とする(1)に記載の方法。 (2) The method according to (1), wherein the storing step stores the number of users for the identification data received by the receiving step during a predetermined period.
このような構成によれば、当該方法を実行するサーバは、所定の期間において識別データの登録数(ユーザ数)を集計する。したがって、当該サーバは、現在の流行に合った索引データを生成することができる。 According to such a configuration, the server that executes the method counts the number of registered identification data (number of users) in a predetermined period. Therefore, the server can generate index data that matches the current trend.
(3) 前記生成ステップは、前記ユーザ数が所定の閾値を超えた場合に、当該ユーザ数に対応する前記組み合わせについて、前記索引データを生成することを特徴とする(1)または(2)に記載の方法。 (3) According to (1) or (2), in the generation step, when the number of users exceeds a predetermined threshold, the index data is generated for the combination corresponding to the number of users. The method described.
このような構成によれば、当該方法を実行するサーバは、識別データの登録数(ユーザ数)が所定の閾値を超えた場合に、この識別データが含むキーワードにより索引データを生成する。したがって、当該サーバは、所定数のユーザからの支持を得られたキーワード集合に含まれる有用性の高いキーワードにより、精度の高い索引データを生成できる可能性がある。 According to such a configuration, when the number of registered identification data (number of users) exceeds a predetermined threshold, the server that executes the method generates index data using the keyword included in the identification data. Therefore, there is a possibility that the server can generate highly accurate index data using highly useful keywords included in the keyword set that has received support from a predetermined number of users.
(4) 前記生成ステップは、前記記憶ステップにより記憶された前記組み合わせの中で、前記ユーザ数が多い順に所定数の組み合わせについて、前記索引データを生成することを特徴とする(1)から(3)のいずれかに記載の方法。 (4) In the generation step, the index data is generated for a predetermined number of combinations in descending order of the number of users among the combinations stored in the storage step. ) Any one of the methods.
このような構成によれば、当該方法を実行するサーバは、識別データの登録数(ユーザ数)が多い順にキーワードを抽出して索引データを生成する。したがって、当該サーバは、現在多くの支持を得ているキーワード集合に含まれる有用性の高いキーワードにより、精度の高い索引データを生成できる可能性がある。 According to such a configuration, the server that executes the method generates index data by extracting keywords in descending order of the number of registered identification data (number of users). Therefore, there is a possibility that the server can generate highly accurate index data by using highly useful keywords included in the keyword set that is currently receiving much support.
(5) 前記ユーザの数に基づいて、前記索引データの優先順位に係る重み付けを調整する調整ステップを更に含む(1)から(4)のいずれかに記載の方法。 (5) The method according to any one of (1) to (4), further including an adjustment step of adjusting a weight related to the priority order of the index data based on the number of users.
このような構成によれば、当該方法を実行するサーバは、識別データの登録数(ユーザ数)に基づいて、例えば、登録数の多い識別データに含まれるキーワードに対するコンテンツの優先順位を高めることができる。その結果、多くのユーザに支持されたキーワードとコンテンツの組み合わせにより、有用性が高いコンテンツを容易に抽出して閲覧できる可能性がある。 According to such a configuration, the server that executes the method can increase the priority order of the content with respect to the keyword included in the identification data having a large number of registrations, for example, based on the number of identification data registered (number of users). it can. As a result, there is a possibility that highly useful content can be easily extracted and browsed by a combination of keywords and content supported by many users.
(6) キーワードと当該キーワードにより検索されるコンテンツとを対応付けた索引データを管理するサーバであって、
ユーザにより当該コンテンツに関連付けられた、識別データを受信する受信手段と、
前記受信手段により受信した識別データに含まれる、所定の種類の語句の集合を抽出する抽出手段と、
前記抽出手段により抽出された語句の集合と、当該語句の集合が抽出された識別データに関連付けられたコンテンツと、の同一の組み合わせについて、当該関連付けが行われたユーザ数を、当該組み合わせと対応付けて記憶する記憶手段と、
前記記憶手段により記憶された前記ユーザ数に基づいて、当該ユーザ数に対応する前記組み合わせについて、前記語句の集合に含まれる語句それぞれを前記キーワードとした前記索引データを生成する生成手段と、を備えるサーバ。
(6) A server that manages index data in which a keyword is associated with content searched by the keyword,
Receiving means for receiving identification data associated with the content by the user;
Extracting means for extracting a set of words of a predetermined type included in the identification data received by the receiving means;
For the same combination of the set of phrases extracted by the extraction unit and the content associated with the identification data from which the set of phrases is extracted, the number of associated users is associated with the combination. Storage means for storing
Generating means for generating the index data using each of the words included in the set of words as the keyword for the combination corresponding to the number of users based on the number of users stored by the storage means; server.
このような構成によれば、当該サーバを運用することにより、(1)と同様の効果が期待できる。 According to such a configuration, the same effect as in (1) can be expected by operating the server.
(7) キーワードと当該キーワードにより検索されるコンテンツとを対応付けた索引データをサーバに管理させるプログラムであって、
ユーザにより当該コンテンツに関連付けられた、識別データを受信する受信ステップと、
前記受信ステップにより受信した識別データに含まれる、所定の種類の語句の集合を抽出する抽出ステップと、
前記抽出ステップにより抽出された語句の集合と、当該語句の集合が抽出された識別データに関連付けられたコンテンツと、の同一の組み合わせについて、当該関連付けが行われたユーザ数を、当該組み合わせと対応付けて記憶する記憶ステップと、
前記記憶ステップにより記憶された前記ユーザ数に基づいて、当該ユーザ数に対応する前記組み合わせについて、前記語句の集合に含まれる語句それぞれを前記キーワードとした前記索引データを生成する生成ステップと、を実行させるプログラム。
(7) A program for causing a server to manage index data in which a keyword is associated with content searched by the keyword,
A receiving step of receiving identification data associated with the content by the user;
An extraction step of extracting a set of words of a predetermined type included in the identification data received by the reception step;
For the same combination of the set of phrases extracted by the extraction step and the content associated with the identification data from which the set of phrases is extracted, the number of users associated is associated with the combination. Memory step for storing
Generating, based on the number of users stored in the storage step, the index data for the combination corresponding to the number of users, using the words included in the set of phrases as the keywords. Program to make.
このような構成によれば、当該プログラムを実行することにより、(1)と同様の効果が期待できる。 According to such a configuration, the same effect as in (1) can be expected by executing the program.
本発明によれば、コンテンツの識別データの有用性に基づく索引データを管理することができる。 According to the present invention, index data based on usefulness of content identification data can be managed.
以下、本発明の実施形態について図を参照しながら説明する。なお、本発明に係るコンテンツをWebページ、識別データをブックマークのタイトルであるとして説明する。 Hereinafter, embodiments of the present invention will be described with reference to the drawings. It is assumed that the content according to the present invention is a Web page and the identification data is a bookmark title.
[システム概要]
図1は、本実施形態に係る検索システムの全体構成を示す概略図である。検索サーバ10は、ブックマーク管理サーバ20、ユーザ端末30、およびコンテンツサーバ40と、ネットワークを介して接続されている。
[System Overview]
FIG. 1 is a schematic diagram showing the overall configuration of the search system according to the present embodiment. The
検索サーバ10は、ユーザ端末30からの要求に応じて、コンテンツサーバ40により管理されるWebページを検索し、ユーザ端末30に検索結果を送信する。
In response to a request from the
ブックマーク管理サーバ20は、ユーザ端末30にて閲覧されたWebページに対するブックマークの登録を受け付ける。そして、ブックマーク管理サーバ20は、ユーザ毎に、ブックマークとWebページのURLとの組み合わせを記憶する。
The
ここで、検索サーバ10は、ブックマーク管理サーバ20から、ブックマーク情報を収集し、登録数に基づく索引データを生成することにより、検索の効率および精度を向上させる。
Here, the
[ハードウェア構成]
図2は、本実施形態に係る検索サーバ10のハードウェア構成の一例を示す図である。検索サーバ10は、制御装置101を構成するCPU(Central Processing Unit)1(1010)(マルチプロセッサ構成ではCPU2(1012)等複数のCPUが追加されてもよい)、バスライン1005、通信I/F1040、メインメモリ1050、BIOS(Basic Input Output System)1060、USBポート1090、I/Oコントローラ1070、ならびにキーボードおよびマウス等の入力装置1100や表示装置1022を備える。
[Hardware configuration]
FIG. 2 is a diagram illustrating an example of a hardware configuration of the
BIOS1060は、検索サーバ10の起動時に制御装置101が実行するブートプログラムや、検索サーバ10のハードウェアに依存するプログラム等を格納する。
The
I/Oコントローラ1070には、テープドライブ1072、ハードディスク1074、光ディスクドライブ1076、半導体メモリ1078等の記憶装置107を接続することができる。
A
記憶装置107を構成するハードディスク1074は、検索サーバ10がサーバとして機能するための各種プログラムおよび本発明の機能を実行するプログラムを記憶しており、更に必要に応じて各種データベースを構成可能である。
The
光ディスクドライブ1076としては、例えば、DVD−ROMドライブ、CD−ROMドライブ、DVD−RAMドライブ、CD−RAMドライブ等を使用することができる。この場合は各ドライブに対応した光ディスク1077を使用する。光ディスク1077から光ディスクドライブ1076によりプログラムまたはデータを読み取り、I/Oコントローラ1070を介してメインメモリ1050またはハードディスク1074に提供することもできる。また、同様にテープドライブ1072に対応したテープメディア1071を主としてバックアップのために使用することもできる。
As the optical disk drive 1076, for example, a DVD-ROM drive, a CD-ROM drive, a DVD-RAM drive, a CD-RAM drive, or the like can be used. In this case, the
検索サーバ10に提供されるプログラムは、ハードディスク1074、光ディスク1077、またはメモリーカード等の記録媒体に格納されて提供される。このプログラムは、I/Oコントローラ1070を介して、記録媒体から読み出され、または通信I/F1040を介してダウンロードされることによって、検索サーバ10にインストールされ実行されてもよい。
The program provided to the
前述のプログラムは、内部または外部の記憶媒体に格納されてもよい。ここで、記憶装置107を構成する記憶媒体としては、ハードディスク1074、光ディスク1077、またはメモリーカードの他に、MD等の光磁気記録媒体、テープ媒体を用いることができる。また、専用通信回線やインターネットに接続されたサーバシステムに設けたハードディスク1074または光ディスクライブラリー等の記憶装置を記録媒体として使用し、通信回線を介してプログラムを検索サーバ10に提供してもよい。
The aforementioned program may be stored in an internal or external storage medium. Here, as a storage medium constituting the
ここで、表示装置1022は、ユーザにデータの入力を受け付ける画面を表示したり、検索サーバ10による演算処理結果の画面を表示したりするものであり、ブラウン管表示装置(CRT)、液晶表示装置(LCD)等のディスプレイ装置を含む。
Here, the
ここで、入力装置1100は、ユーザによる入力の受け付けを行うものであり、キーボードおよびマウス等により構成してよい。
Here, the
また、通信I/F1040は、検索サーバ10を専用ネットワークまたは公共ネットワークを介して端末と接続できるようにするためのネットワーク・アダプタである。通信I/F1040は、モデム、ケーブル・モデムおよびイーサネット(登録商標)・アダプタを含んでよい。
The communication I /
以上の例は、検索サーバ10について主に説明したが、コンピュータに、プログラムをインストールして、そのコンピュータをサーバ装置として動作させることにより上記で説明した機能を実現することもできる。したがって、本発明において一実施形態として説明した検索サーバ10により実現される機能は、上述の方法を当該コンピュータで実行することにより、あるいは、上述のプログラムを当該コンピュータに導入して実行することによっても実現可能である。
In the above example, the
[機能構成]
図3は、本実施形態に係る検索サーバ10の機能構成を示すブロック図である。検索サーバ10は、制御装置101に、ブックマーク受信部11と、キーワード集合抽出部12と、ブックマークカウント部13と、索引生成部14と、優先度調整部15と、を備え、記憶装置107に、カウントDB16と、索引DB17と、を備える。
[Function configuration]
FIG. 3 is a block diagram showing a functional configuration of the
ブックマーク受信部11は、ブックマーク管理サーバ20が備えるブックマークDB21から、ブックマークデータを受信する。ここで、ブックマークデータは、ユーザ端末30のユーザにより登録されたWebページの識別データ(ブックマークのタイトル)とWebページのURLとを対応付けたデータである。
The
図4は、本実施形態に係るブックマークテーブルを示す図である。ここでは、ブックマークを登録したユーザのIDと、ブックマークのタイトルと、WebページのURLと、ブックマークの登録日時と、が関連付けて記憶されている。各ユーザは、自分が登録したブックマークのタイトルを選択することにより、対応するURLのWebページを取得し閲覧することができる。また、ブックマークは、他者へ公開できてもよく、ブックマークテーブルには、公開するか否かのフラグが付与される。 FIG. 4 is a diagram showing a bookmark table according to the present embodiment. Here, the ID of the user who registered the bookmark, the bookmark title, the URL of the web page, and the bookmark registration date and time are stored in association with each other. Each user can acquire and browse a Web page with a corresponding URL by selecting a bookmark title registered by the user. In addition, the bookmark may be open to others, and a flag indicating whether to open the bookmark is given to the bookmark table.
ブックマーク受信部11は、ブックマークテーブル(図4)に記憶されたブックマークデータのうち、登録日時が所定の期間のデータを受信する。この受信は所定の周期で行うこととしてよく、前回の受信から現在までに登録されたブックマークデータを受信する。
The
なお、受信するブックマークデータは、公開されたものに限定してもよい。このことによれば、検索サーバ10は、ユーザが他者に見せることを前提に名付けたブックマークのタイトルを利用して索引データを生成する。したがって、他者にも支持され易い有用なキーワードによる索引データが生成され、検索の精度を向上させられる可能性がある。
Note that the received bookmark data may be limited to published data. According to this, the
キーワード集合抽出部12は、ブックマーク受信部11により受信したブックマークのタイトルに含まれる、キーワードの集合を抽出する。具体的には、キーワード集合抽出部12は、ブックマークのタイトルを形態素解析することにより、名詞等、所定の種類の単語を抽出し、抽出した単語の組み合わせをキーワード集合とする。
The keyword set
ブックマークカウント部13は、キーワード集合抽出部12により抽出されたキーワード集合と、対応するブックマークデータに含まれるURLと、の組み合わせが現れた回数をカウントし、カウントDB16に記憶する。すなわち、ブックマークカウント部13は、同一のキーワード集合が含まれた同様のブックマークのタイトルが同一のWebページに対して登録された場合に、その登録数をカウントして記憶する。
The
図5は、本実施形態に係るカウントDB16に記憶されるカウントテーブルを示す図である。ここでは、キーワードの集合とWebページのURLとの組み合わせに対して、その組み合わせのブックマークを登録したユーザ数が記憶される。
FIG. 5 is a diagram showing a count table stored in the
索引生成部14は、カウントテーブル(図5)を参照し、検索処理に用いられる索引データを生成する。ここで、索引データは、キーワードとURLとの組み合わせを含むデータであって、Webページの検索処理において、検索エンジンが参照する。
The
具体的には、索引生成部14は、カウントテーブル(図5)から、登録数が相対的に多いキーワード集合とURLの組み合わせを抽出する。例えば、登録数が所定の閾値を超えている組み合わせや、登録数が多い順に所定数の組み合わせを抽出することとしてよい。そして、索引生成部14は、抽出した組み合わせについて、キーワード集合が含む個々のキーワードそれぞれとURLとを対応付けた索引データを生成し、索引DB17に記憶する。
Specifically, the
なお、索引データは、個々のキーワードとURLとを対応付けることとしたが、これには限られず、キーワード集合または部分集合とURLとを対応付けてもよい。 Although the index data associates each keyword with a URL, the present invention is not limited to this, and a keyword set or a subset may be associated with a URL.
優先度調整部15は、カウントテーブル(図5)の登録数に基づいて、索引DB17に記憶された索引データに関して、検索の優先順位を示す優先度の調整を行う。優先度は、検索結果としてのURLの順位(表示順序)を決定するための重み付けであってよく、登録数が多いほど、優先度を高く調整する。これにより、多数のユーザのブックマークで対応付けがなされたキーワードとURLの組み合わせについては、優先度が高まるため、検索され易くなる。
The
以上で説明した各機能は、検索サーバ10の各部により実現されることとしたが、これには限られず、適宜、複数のサーバに分散させてもよい。
Each function described above is realized by each unit of the
[第1の処理]
図6は、本実施形態に係る検索サーバ10における第1の処理の流れを示す図である。この処理では、キーワード集合とURLとの組み合わせの登録数が所定の閾値以上である場合に索引データを生成する。
[First processing]
FIG. 6 is a diagram showing a flow of the first process in the
ステップS1では、制御装置101は、カウントDB16のカウントテーブル(図5)から、キーワード集合とURLの組み合わせを1件抽出する。
In step S1, the
ステップS2では、制御装置101は、ステップS1にて抽出したデータについて、登録数が所定の閾値以上であるか否かを判定する。この判定がYESの場合はステップS3に移り、判定がNOの場合はステップS6に移る。
In step S2, the
ステップS3では、制御装置101は、ステップS2にて登録数が閾値以上であると判定されたキーワード集合とURLとの組み合わせに関して、索引データが既に登録済みであるか否かを判定する。この判定がYESの場合はステップS5に移り、判定がNOの場合はステップS4に移る。
In step S3, the
ステップS4では、制御装置101は、ステップS3にて索引データが登録されていないと判定されたキーワードとURLの組み合わせに関して、索引データを生成して索引DB17に登録する。
In step S4, the
ステップS5では、制御装置101は、登録された索引データに対して、カウントテーブル(図5)から抽出した登録数に基づいて、検索の優先順位に関わる重み付けを行う。
In step S5, the
ステップS6では、制御装置101は、カウントテーブル(図5)に記憶された全件を処理したか否かを判定する。この判定がYESの場合は処理を終了し、判定がNOの場合はステップS1に戻って次の1件に対して処理を行う。
In step S6, the
[第2の処理]
図7は、本実施形態に係る検索サーバ10における第2の処理の流れを示す図である。この処理では、キーワード集合とURLの組み合わせのうち、登録数が最も多いもの、あるいは登録数が多いものから所定数について、索引データを生成する。
[Second processing]
FIG. 7 is a diagram showing a flow of second processing in the
ステップS11では、制御装置101は、カウントテーブル(図5)のデータを、登録数によりソートする。
In step S11, the
ステップS12では、制御装置101は、ステップS11にてソートしたデータのうち、登録数が上位のものから順に、キーワード集合とURLの組み合わせを1件抽出する。
In step S12, the
ステップS13では、制御装置101は、ステップS12にて抽出したデータについて、登録数が所定の閾値以上であるか否かを判定する。この判定がYESの場合はステップS14に移り、判定がNOの場合は、登録数が閾値以上のデータがこれ以上無いと判断して、処理を終了する。
In step S13, the
ステップS14では、制御装置101は、ステップS13にて登録数が閾値以上であると判定されたキーワード集合とURLとの組み合わせに関して、索引データが既に登録済みであるか否かを判定する。この判定がYESの場合はステップS16に移り、判定がNOの場合はステップS15に移る。
In step S14, the
ステップS15では、制御装置101は、ステップS14にて索引データが登録されていないと判定されたキーワードとURLの組み合わせに関して、索引データを生成して索引DB17に登録する。
In step S15, the
ステップS16では、制御装置101は、登録された索引データに対して、カウントテーブル(図5)から抽出した登録数に基づいて、検索の優先順位に関わる重み付けを行う。
In step S16, the
ステップS17では、制御装置101は、カウントテーブル(図5)から索引データの生成対象となる所定数を抽出したか否かを判定する。この判定がYESの場合は処理を終了し、判定がNOの場合はステップS12に戻って次の1件に対して処理を行う。
In step S17, the
以上、本発明の実施形態について説明したが、本発明は上述した実施形態に限るものではない。また、本発明の実施形態に記載された効果は、本発明から生じる最も好適な効果を列挙したに過ぎず、本発明による効果は、本発明の実施形態に記載されたものに限定されるものではない。 As mentioned above, although embodiment of this invention was described, this invention is not restricted to embodiment mentioned above. The effects described in the embodiments of the present invention are only the most preferable effects resulting from the present invention, and the effects of the present invention are limited to those described in the embodiments of the present invention. is not.
10 検索サーバ
11 ブックマーク受信部
12 キーワード集合抽出部
13 ブックマークカウント部
14 索引生成部
15 優先度調整部
16 カウントDB
17 索引DB
20 ブックマーク管理サーバ
21 ブックマークDB
30 ユーザ端末
40 コンテンツサーバ
101 制御装置
107 記憶装置
DESCRIPTION OF
17 Index DB
20
30
Claims (7)
ユーザにより当該コンテンツに関連付けられた、識別データを受信する受信ステップと、
前記受信ステップにより受信した識別データに含まれる、所定の種類の語句の集合を抽出する抽出ステップと、
前記抽出ステップにより抽出された語句の集合と、当該語句の集合が抽出された識別データに関連付けられたコンテンツと、の同一の組み合わせについて、当該関連付けが行われたユーザ数を、当該組み合わせと対応付けて記憶する記憶ステップと、
前記記憶ステップにより記憶された前記ユーザ数に基づいて、当該ユーザ数に対応する前記組み合わせについて、前記語句の集合に含まれる語句それぞれを前記キーワードとした前記索引データを生成する生成ステップと、を含む方法。 A server manages index data in which a keyword is associated with content searched for by the keyword,
A receiving step of receiving identification data associated with the content by the user;
An extraction step of extracting a set of words of a predetermined type included in the identification data received by the reception step;
For the same combination of the set of phrases extracted by the extraction step and the content associated with the identification data from which the set of phrases is extracted, the number of users associated is associated with the combination. Memory step for storing
Generating based on the number of users stored in the storing step, the index data for the combination corresponding to the number of users, the index data using each of the words included in the set of words as the keyword. Method.
ユーザにより当該コンテンツに関連付けられた、識別データを受信する受信手段と、
前記受信手段により受信した識別データに含まれる、所定の種類の語句の集合を抽出する抽出手段と、
前記抽出手段により抽出された語句の集合と、当該語句の集合が抽出された識別データに関連付けられたコンテンツと、の同一の組み合わせについて、当該関連付けが行われたユーザ数を、当該組み合わせと対応付けて記憶する記憶手段と、
前記記憶手段により記憶された前記ユーザ数に基づいて、当該ユーザ数に対応する前記組み合わせについて、前記語句の集合に含まれる語句それぞれを前記キーワードとした前記索引データを生成する生成手段と、を備えるサーバ。 A server that manages index data in which a keyword is associated with content searched by the keyword,
Receiving means for receiving identification data associated with the content by the user;
Extracting means for extracting a set of words of a predetermined type included in the identification data received by the receiving means;
For the same combination of the set of phrases extracted by the extraction unit and the content associated with the identification data from which the set of phrases is extracted, the number of associated users is associated with the combination. Storage means for storing
Generating means for generating the index data using each of the words included in the set of words as the keyword for the combination corresponding to the number of users based on the number of users stored by the storage means; server.
ユーザにより当該コンテンツに関連付けられた、識別データを受信する受信ステップと、
前記受信ステップにより受信した識別データに含まれる、所定の種類の語句の集合を抽出する抽出ステップと、
前記抽出ステップにより抽出された語句の集合と、当該語句の集合が抽出された識別データに関連付けられたコンテンツと、の同一の組み合わせについて、当該関連付けが行われたユーザ数を、当該組み合わせと対応付けて記憶する記憶ステップと、
前記記憶ステップにより記憶された前記ユーザ数に基づいて、当該ユーザ数に対応する前記組み合わせについて、前記語句の集合に含まれる語句それぞれを前記キーワードとした前記索引データを生成する生成ステップと、を実行させるプログラム。 A program that causes a server to manage index data that associates a keyword with content searched for by the keyword,
A receiving step of receiving identification data associated with the content by the user;
An extraction step of extracting a set of words of a predetermined type included in the identification data received by the reception step;
For the same combination of the set of phrases extracted by the extraction step and the content associated with the identification data from which the set of phrases is extracted, the number of users associated is associated with the combination. Memory step for storing
Generating, based on the number of users stored in the storage step, the index data for the combination corresponding to the number of users, using the words included in the set of phrases as the keywords. Program to make.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2008131603A JP5014252B2 (en) | 2008-05-20 | 2008-05-20 | Method, server, and program for managing index data for searching content |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2008131603A JP5014252B2 (en) | 2008-05-20 | 2008-05-20 | Method, server, and program for managing index data for searching content |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JP2009282593A true JP2009282593A (en) | 2009-12-03 |
| JP5014252B2 JP5014252B2 (en) | 2012-08-29 |
Family
ID=41453022
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2008131603A Active JP5014252B2 (en) | 2008-05-20 | 2008-05-20 | Method, server, and program for managing index data for searching content |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JP5014252B2 (en) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2014071645A (en) * | 2012-09-28 | 2014-04-21 | Ntt Docomo Inc | Server device, information processing method and program |
| JP2018055605A (en) * | 2016-09-30 | 2018-04-05 | ジャパンモード株式会社 | Innovation creation support program |
| JP2018055604A (en) * | 2016-09-30 | 2018-04-05 | ジャパンモード株式会社 | Creation support program |
| JP2021108124A (en) * | 2019-12-29 | 2021-07-29 | Deiシステムズ株式会社 | Access target retrieval system |
Citations (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH10162126A (en) * | 1996-12-02 | 1998-06-19 | Planet Computer:Kk | Electronization device for document |
| JP2002014996A (en) * | 2000-06-30 | 2002-01-18 | Nec Corp | Bookmark system, document proposal method using bookmark and program recording medium |
| JP2003067328A (en) * | 2001-08-29 | 2003-03-07 | Nec Corp | Bookmark management system and bookmark management method |
| JP2003271670A (en) * | 2002-03-19 | 2003-09-26 | Mitsubishi Electric Corp | Information collecting apparatus, information collecting method and program |
| JP2005010882A (en) * | 2003-06-17 | 2005-01-13 | Hidekazu Kondo | Browser device, information processing system and program |
| WO2005008527A1 (en) * | 2003-07-16 | 2005-01-27 | Fujitsu Limited | Dynamically categorized bookmark management device |
| JP2006107200A (en) * | 2004-10-06 | 2006-04-20 | Vodafone Kk | Search service provision system |
| WO2006095409A1 (en) * | 2005-03-07 | 2006-09-14 | Mars Flag Corporation | Information retrieving device, computer program, and recording medium |
| JP2007072596A (en) * | 2005-09-05 | 2007-03-22 | Nippon Telegr & Teleph Corp <Ntt> | Information sharing system and information sharing method |
| JP2007133794A (en) * | 2005-11-14 | 2007-05-31 | Hitachi Ltd | Electronic document management apparatus, electronic document management program, electronic document management system |
| JP2008112355A (en) * | 2006-10-31 | 2008-05-15 | Fujitsu Ltd | Bookmark management device, bookmark management program, and bookmark management method |
-
2008
- 2008-05-20 JP JP2008131603A patent/JP5014252B2/en active Active
Patent Citations (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH10162126A (en) * | 1996-12-02 | 1998-06-19 | Planet Computer:Kk | Electronization device for document |
| JP2002014996A (en) * | 2000-06-30 | 2002-01-18 | Nec Corp | Bookmark system, document proposal method using bookmark and program recording medium |
| JP2003067328A (en) * | 2001-08-29 | 2003-03-07 | Nec Corp | Bookmark management system and bookmark management method |
| JP2003271670A (en) * | 2002-03-19 | 2003-09-26 | Mitsubishi Electric Corp | Information collecting apparatus, information collecting method and program |
| JP2005010882A (en) * | 2003-06-17 | 2005-01-13 | Hidekazu Kondo | Browser device, information processing system and program |
| WO2005008527A1 (en) * | 2003-07-16 | 2005-01-27 | Fujitsu Limited | Dynamically categorized bookmark management device |
| JP2006107200A (en) * | 2004-10-06 | 2006-04-20 | Vodafone Kk | Search service provision system |
| WO2006095409A1 (en) * | 2005-03-07 | 2006-09-14 | Mars Flag Corporation | Information retrieving device, computer program, and recording medium |
| JP2007072596A (en) * | 2005-09-05 | 2007-03-22 | Nippon Telegr & Teleph Corp <Ntt> | Information sharing system and information sharing method |
| JP2007133794A (en) * | 2005-11-14 | 2007-05-31 | Hitachi Ltd | Electronic document management apparatus, electronic document management program, electronic document management system |
| JP2008112355A (en) * | 2006-10-31 | 2008-05-15 | Fujitsu Ltd | Bookmark management device, bookmark management program, and bookmark management method |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2014071645A (en) * | 2012-09-28 | 2014-04-21 | Ntt Docomo Inc | Server device, information processing method and program |
| JP2018055605A (en) * | 2016-09-30 | 2018-04-05 | ジャパンモード株式会社 | Innovation creation support program |
| JP2018055604A (en) * | 2016-09-30 | 2018-04-05 | ジャパンモード株式会社 | Creation support program |
| JP2021108124A (en) * | 2019-12-29 | 2021-07-29 | Deiシステムズ株式会社 | Access target retrieval system |
| JP7157474B2 (en) | 2019-12-29 | 2022-10-20 | Deiシステムズ株式会社 | Access target search system |
Also Published As
| Publication number | Publication date |
|---|---|
| JP5014252B2 (en) | 2012-08-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP4342575B2 (en) | Device, method, and program for keyword presentation | |
| CA2647864C (en) | Propagating useful information among related web pages, such as web pages of a website | |
| JP5264892B2 (en) | Multilingual information search | |
| US9183261B2 (en) | Lexicon based systems and methods for intelligent media search | |
| US20080294619A1 (en) | System and method for automatic generation of search suggestions based on recent operator behavior | |
| EP1887485A2 (en) | Keyword outputting apparatus, keyword outputting method, and keyword outputting computer program product | |
| JP4962986B2 (en) | Method, server, and program for classifying content data into categories | |
| US7310633B1 (en) | Methods and systems for generating textual information | |
| JP4934169B2 (en) | Apparatus, method, and program for associating categories | |
| US9754022B2 (en) | System and method for language sensitive contextual searching | |
| WO2012012396A2 (en) | Predictive query suggestion caching | |
| JP5226241B2 (en) | How to add tags | |
| US20160217181A1 (en) | Annotating Query Suggestions With Descriptions | |
| US20150339387A1 (en) | Method of and system for furnishing a user of a client device with a network resource | |
| JP5014252B2 (en) | Method, server, and program for managing index data for searching content | |
| JP4962945B2 (en) | Bookmark / tag setting device | |
| KR100455439B1 (en) | Internet resource retrieval and browsing method based on expanded web site map and expanded natural domain names assigned to all web resources | |
| JP4796527B2 (en) | Document narrowing search apparatus, method and program | |
| JP4819628B2 (en) | Method, server, and program for retrieving document data | |
| JP5072792B2 (en) | Retrieval method, program and server for preferentially displaying pages according to amount of information | |
| JP5285491B2 (en) | Information retrieval system, method and program, index creation system, method and program, | |
| CN107818091B (en) | Document processing method and device | |
| JP6079207B2 (en) | Keyword presentation program, keyword presentation method, and keyword presentation apparatus | |
| JP2008262442A (en) | Method and server for displaying search key data | |
| JP5416023B2 (en) | Reading terminal and method |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| A977 | Report on retrieval |
Free format text: JAPANESE INTERMEDIATE CODE: A971007 Effective date: 20111031 |
|
| A131 | Notification of reasons for refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A131 Effective date: 20111108 |
|
| A521 | Request for written amendment filed |
Free format text: JAPANESE INTERMEDIATE CODE: A523 Effective date: 20120106 |
|
| RD04 | Notification of resignation of power of attorney |
Free format text: JAPANESE INTERMEDIATE CODE: A7424 Effective date: 20120312 |
|
| TRDD | Decision of grant or rejection written | ||
| A01 | Written decision to grant a patent or to grant a registration (utility model) |
Free format text: JAPANESE INTERMEDIATE CODE: A01 Effective date: 20120515 |
|
| A01 | Written decision to grant a patent or to grant a registration (utility model) |
Free format text: JAPANESE INTERMEDIATE CODE: A01 |
|
| A61 | First payment of annual fees (during grant procedure) |
Free format text: JAPANESE INTERMEDIATE CODE: A61 Effective date: 20120605 |
|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20150615 Year of fee payment: 3 |
|
| R150 | Certificate of patent or registration of utility model |
Free format text: JAPANESE INTERMEDIATE CODE: R150 Ref document number: 5014252 Country of ref document: JP Free format text: JAPANESE INTERMEDIATE CODE: R150 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| S531 | Written request for registration of change of domicile |
Free format text: JAPANESE INTERMEDIATE CODE: R313531 |
|
| R350 | Written notification of registration of transfer |
Free format text: JAPANESE INTERMEDIATE CODE: R350 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| S533 | Written request for registration of change of name |
Free format text: JAPANESE INTERMEDIATE CODE: R313533 |
|
| R350 | Written notification of registration of transfer |
Free format text: JAPANESE INTERMEDIATE CODE: R350 |
|
| S111 | Request for change of ownership or part of ownership |
Free format text: JAPANESE INTERMEDIATE CODE: R313111 |
|
| R350 | Written notification of registration of transfer |
Free format text: JAPANESE INTERMEDIATE CODE: R350 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| S111 | Request for change of ownership or part of ownership |
Free format text: JAPANESE INTERMEDIATE CODE: R313111 |
|
| R350 | Written notification of registration of transfer |
Free format text: JAPANESE INTERMEDIATE CODE: R350 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |