About: SpiderLing

Facets (new session)
Description
Metadata
Settings
- owl:sameAs
- Inference Rule:

About: SpiderLing Goto Sponge NotDistinct Permalink

An Entity of Type : http://linked.opendata.cz/ontology/domain/vavai/Vysledek, within Data Space : linked.opendata.cz associated with source document(s)

Attributes	Values
rdf:type	skos:Concept http://linked.opendata.cz/ontology/domain/vavai/Vysledek
rdfs:seeAlso	http://nlp.fi.muni.cz/trac/spiderling/
Description	SpiderLing - a web spider for linguistics - is a software for obtaining large textual data from the web. The purpose of the obtained data is building text corpora. Many documents on the web only contain material not suitable for text corpora, such as site navigation, lists of links, lists of products, and other kind of text not comprised of full sentences. In fact such pages represent the vast majority of the web. Therefore, by doing unrestricted web crawls, we typically download a lot of data which gets filtered out during post-processing. This makes the process of web corpus collection inefficient. SpiderLing focuses the crawling on the text rich parts of the web and maximizes the number of words in the final corpus per downloaded megabyte. The crawler was used for building large corpora in American Spanish, Arabic, Czech, English, Estonian, French, Hungarian, Japanese, Korean, Polish, Russian, Tajik, and six Turkic languages consisting of 74 billion words altogether. SpiderLing - a web spider for linguistics - is a software for obtaining large textual data from the web. The purpose of the obtained data is building text corpora. Many documents on the web only contain material not suitable for text corpora, such as site navigation, lists of links, lists of products, and other kind of text not comprised of full sentences. In fact such pages represent the vast majority of the web. Therefore, by doing unrestricted web crawls, we typically download a lot of data which gets filtered out during post-processing. This makes the process of web corpus collection inefficient. SpiderLing focuses the crawling on the text rich parts of the web and maximizes the number of words in the final corpus per downloaded megabyte. The crawler was used for building large corpora in American Spanish, Arabic, Czech, English, Estonian, French, Hungarian, Japanese, Korean, Polish, Russian, Tajik, and six Turkic languages consisting of 74 billion words altogether. (en)
Title	SpiderLing SpiderLing (en)
skos:prefLabel	SpiderLing SpiderLing (en)
skos:notation	RIV/00216224:14330/12:00064706!RIV13-MSM-14330___
http://linked.open...avai/riv/aktivita	P S
http://linked.open...avai/riv/aktivity	P(LM2010013), S
http://linked.open...vai/riv/dodaniDat	2013
http://linked.open...aciTvurceVysledku	Suchomel, Vít
http://linked.open.../riv/druhVysledku	R - Software
http://linked.open...iv/duvernostUdaju	S - Úplné a pravdivé údaje nepodléhající ochraně podle zvláštních právních předpisů
http://linked.open...onomickeParametry	Úspory: Oproti dříve používanému nástroji Heritrix bylo - sníženo zatížení univerzitních strojů a sítě (zkrácen strojový čas potřebný k běhu programu a sníženo množství přenesených dat) - dosaženo vyšší výtěžnosti získaných dat Umožnění navazujících projektů: V případě nasazení méně specializovaného nástroje (běžného crawleru) by bylo nutno využít více prostředků po delší dobu, tudíž by nebylo sestaveno takové množství textových korpusů a nemohly úspěšně navázat projekty uživatelů korpusů (např. vývoj morfologického analyzátoru tádžické perštiny, výuka angličtiny, jazykové analýzy a další). Velké textové korpusy sestavené v Centru zpracování přirozeného jazyka na Fakultě informatiky Masarykovy univerzity jsou zdarma dostupné pro zaměstance a studenty Masarykovy univerzity. Některé korpusy jsou využívány také Oddělením anglistiky Filozofické fakulty Univerzity Palackého Olomouc.
http://linked.open...titaPredkladatele	Masarykova univerzita / Fakulta informatiky
http://linked.open...dnocenehoVysledku	170426
http://linked.open...ai/riv/idVysledku	RIV/00216224:14330/12:00064706
http://linked.open...terniIdentifikace	SpiderLing
http://linked.open...riv/jazykVysledku	eng - angličtina
http://linked.open.../riv/klicovaSlova	web crawler; web spider; text corpora (en)
http://linked.open.../riv/klicoveSlovo	web crawler web spider text corpora
http://linked.open...ontrolniKodProRIV	[F1A937F6A7F4]
http://linked.open.../licencniPoplatek	N - Poskytovatel licence na výsledek nepožaduje licenční poplatek
http://linked.open...in/vavai/riv/obor	AI
http://linked.open...ichTvurcuVysledku	1 (xsd:int)
http://linked.open...cetTvurcuVysledku	1 (xsd:int)
http://linked.open...vavai/riv/projekt	LINDAT-CLARIN: Institute for analysis, processing and distribution of linguistic data
http://linked.open...UplatneniVysledku	2012
http://linked.open...echnickeParametry	Software pro získávání velkého množství textových dokumentů z internetu. Implementace v jazyce Python. Licence: GNU General Public license. Odpovědná osoba pro jednání: Mgr. Pavel Rychlý, Ph.D.; email: pary@fi.muni.cz; telefon: 549496399; adresa: Pavel Rychlý, Fakulta informatiky Masarykovy univerzity, Botanická 68a, 602 00 Brno.
http://linked.open...iv/tvurceVysledku	Suchomel, Vít
http://linked.open...avai/riv/vlastnik	Masarykova univerzita
http://linked.open...itiJinymSubjektem	A - Nabytí licence je nutné vždy
http://localhost/t...ganizacniJednotka	14330

Faceted Search & Find service v1.16.118 as of Jun 21 2024

Alternative Linked Data Documents: ODE Content Formats:

RDF

ODATA

Microdata

About

OpenLink Virtuoso version 07.20.3240 as of Jun 21 2024, on Linux (x86_64-pc-linux-gnu), Single-Server Edition (126 GB total memory, 110 GB memory in use)
Data on this page belongs to its respective rights holders.
Virtuoso Faceted Browser Copyright © 2009-2024 OpenLink Software