Friday, November 26, 2010
Muddiest Point 11/22/10
In class we discussed the z39.50 search. I also remember seeing this in our Koha project assignment. In the slides it stated that in general these types of searches are not used due to the fact that they are difficult to implement, can have semantic problems, and network server problems. So then who really uses it? And why would they use it over another type of search, especially since it does not retrieve full documents or other objects?
Saturday, November 20, 2010
Reading Notes for 11/22/10
1.) I really liked the David Hawking articles about the basics of the web crawlers, how they function, the problems they have, and why they are such great tools. How they must leave out "low-value automated content" since that is not necessary and makes the search harder. How they must be able to search in multiple different languages, misspellings, and also made up words (ie: google, yahoo). How they must work with such a vast amount of information, how they must avoid tricky spamming sites, and all within a few milliseconds. Its quite amazing when you think about it like that!
2.) The Shreeves, S. L et al. article had an example of its customers that I thought was intriguing. The Sheet Music Consortium is trying to digitize sheet music and they initially had some trouble with figuring out how to digitize the various components of their information such as the "cover art, the sheet music itself, the lyrics, etc." and also I assume all the other small additions written onto the music such as playing in piano, forte, or staccato. In class, sometimes we discuss the concept of using other languages for digitizing data, but music was not something I had though of before. Essentially music is like another language.
3.) The Bergman article was great because it explained the true depth of the deep web and the difficulty for crawlers to find this information. Even though the crawlers do a great job with all that they are responsible for, there is still a massive amount of information which is left out of the equation. The article discusses how "Internet searchers are therefore searching only 0.03% — or one in 3,000 — of the pages available to them today." That is such a tiny fraction of the information that we could be accessing! I find that number almost unbelievable. The article also discusses, that crawlers aren't even looking in fire-walled or "Intranet" sites within institutions as they can't. So there's more information there that we cannot access..
2.) The Shreeves, S. L et al. article had an example of its customers that I thought was intriguing. The Sheet Music Consortium is trying to digitize sheet music and they initially had some trouble with figuring out how to digitize the various components of their information such as the "cover art, the sheet music itself, the lyrics, etc." and also I assume all the other small additions written onto the music such as playing in piano, forte, or staccato. In class, sometimes we discuss the concept of using other languages for digitizing data, but music was not something I had though of before. Essentially music is like another language.
3.) The Bergman article was great because it explained the true depth of the deep web and the difficulty for crawlers to find this information. Even though the crawlers do a great job with all that they are responsible for, there is still a massive amount of information which is left out of the equation. The article discusses how "Internet searchers are therefore searching only 0.03% — or one in 3,000 — of the pages available to them today." That is such a tiny fraction of the information that we could be accessing! I find that number almost unbelievable. The article also discusses, that crawlers aren't even looking in fire-walled or "Intranet" sites within institutions as they can't. So there's more information there that we cannot access..
Friday, November 19, 2010
Saturday, November 13, 2010
Reading Notes for 11/15
In the Mischo article it was interesting to read about the university projects that were funded by the DL-1 for networking and computing technologies. It was good to see Carnegie Mellon listed there as I'm sure these grants were only for the most qualified and capable groups. It said that CMU had a grant for the "study of integrated speech, image, video, and language understanding software under its Informedia system." I wonder if that is for speech recognition technology? Or if it was something that was specific for individuals who are either deaf, hearing impaired, or blind?
The Paepcke, A. et al. article is interesting because he talks about "the binary union between academic librarians and computer scientists." And I feel in many ways thats what Information Science is about, a fusion of computer science and library science. I didn't think I really understood till I started my MLIS, the interconnected weave that technology and information have with each other. I already feel very knowledgeable about many computer based concepts. Granted, I have a long long ways to go till I can really feel proficient. But, I still feel an understanding of the concept of this union.
The article by Lynch, Clifford was a very heavy read. But, as I was reading I started wondering about the University of Pittsburgh's Institutional repository. I wondered if we had one, what it is called and who uses it. At first I thought it may be something like Blackboard. But, then it described it more as a place that individuals can access anything by a professor or a graduate student. And on Blackboard you can only access the classes that you have or teach. Therefore, I don't think it fits the description. So I guess I would like to know more about the University of Pittsburgh and institutional repositories.
The Paepcke, A. et al. article is interesting because he talks about "the binary union between academic librarians and computer scientists." And I feel in many ways thats what Information Science is about, a fusion of computer science and library science. I didn't think I really understood till I started my MLIS, the interconnected weave that technology and information have with each other. I already feel very knowledgeable about many computer based concepts. Granted, I have a long long ways to go till I can really feel proficient. But, I still feel an understanding of the concept of this union.
The article by Lynch, Clifford was a very heavy read. But, as I was reading I started wondering about the University of Pittsburgh's Institutional repository. I wondered if we had one, what it is called and who uses it. At first I thought it may be something like Blackboard. But, then it described it more as a place that individuals can access anything by a professor or a graduate student. And on Blackboard you can only access the classes that you have or teach. Therefore, I don't think it fits the description. So I guess I would like to know more about the University of Pittsburgh and institutional repositories.
Subscribe to:
Posts (Atom)