|By Ben Forta||
|January 9, 2002 12:00 AM EST||
Barely a week goes by without someone asking me about ColdFusion and search engine-friendly URLs. This is one of those topics that ColdFusion developers have been discussing for a long time - I first started a thread on this subject on the Allaire Developer's Forum close to five years ago. As this topic keeps coming up (and because three of you e-mailed me to ask about it this morning), I decided to scrap the column I was writing in favor of an explanation of all this once and for all.
Search Engines and ColdFusion
Here's the problem. Many search engines index sites by spidering them - starting at a known point, indexing that retrieved page, parsing it for embedded URLs, and then retrieving and indexing them too. They keep doing this until they have retrieved and indexed all linked pages in a site.
At least that's how it's supposed to work. Where things get complicated is when dynamic pages are indexed, or rather, when they're not. Some spiders won't index URLs that are dynamic, the theory being that since the content changes all the time there's no point in indexing it - after all, it may not even be the same content on subsequent accesses. So this link would be followed and indexed:
but this one may not:
You can see the problem: ColdFusion developers (as well as ASP, PHP, Perl, Python, and JSP developers, among others) build dynamic pages. That's why we use ColdFusion - if all content was static we'd use plain old HTML without any server-side processing at all. We use ColdFusion because we want and need dynamic content - content that spiders may choose to ignore. If search engine indexing is a requirement for you, that can pose a real problem.
The Basic Solution
The solution to this problem is actually quite simple. What is it that tips off the spider, telling it that the page is dynamic? It's the query string portion of the URL, starting with the question mark. That question mark separates the URL (the file to be retrieved or the script to be executed) from any passed parameters (usually in name=value pairs).
The solution is to simply not have a query string - no question mark, no values passed after it, and no name=value pairs. Simple, eh? Yep, until you actually need passed parameters. Then what?
Well, look at this URL:
What happens when this is processed? catalog.cfm is the file to be executed and anything after the file name is ignored altogether. Even though the Web server and ColdFusion ignore it, that doesn't mean you have to. Using CGI variables (like PATH_INFO) you can access the complete URL - even parts of it that may be ignored by other processes or applications. The following snippet will display the complete path (minus any host and protocol information, although that's available in another CGI variable if needed).
Using the above URL this will display:
We now know how to get the passed information. Next you need to remove the path and script name. You could do this by searching for the CFM file, but there's an easier way using yet another CGI variable, this time SCRIPT_NAME, which contains the name of the script being executed (here /path/catalog.cfm):
<CFSET query_string=Right(CGI.PATH_INFO, query_string_length)>
The first <CFSET> gets the length of PATH_INFO and subtracts the length of SCRIPT_NAME from it, giving us the length of the section at the end (the data we want). The second <CFSET> uses Right() to extract the desired data, saving it in a variable named query_string. Displaying query_string would result in this output:
That's how you get the data off the end of the URL.
Note: Depending on the Web server and OS being used you may find different CGI variables or different values in them. You can use <CFDUMP VAR="#CGI#"> to see all the CGI variables available to you (and you'll be able to adapt the code here accordingly).
Making It All Work
Now that the basic concept is clear, let's take this one step further. How can you pass multiple parameters? Well, you could pass three parameters, each separated by a /:
But you would have no way of knowing which was which, what their names should be, and they'd always have to be in order (something you typically don't worry about when working with URL parameters).
While the technique is sound, the implementation can be improved by simulating name=value pairs. Look at this example:
There are now six values passed to the URL; each set of two is a name and value pair. Here it's dept=10, user=0, and item=B12. The advantage of this enhancement is that the variable name is known, the order of pairs is irrelevant, and the URL is easy to construct.
All that is required now is a simple way to extract these values and turn them into real URL values (after all, your application code shouldn't have to worry about manipulating this data; as far as it's concerned these are URL parameters plain and simple). This is where ColdFusion's list functions come in handy:
<!--- How many items in "query_string? --->
<!--- Loop through list, pair of items at a time --->
<!--- Extract "query_string" from full path --->
The first two <CFSET> statements simply extract the virtual query_string (as we saw earlier). Next we need to know how many items query_string contains. All ColdFusion list functions take an optional delimiter as a parameter, and so ListLen(query_string, "/") returns the number of items delimited by / (in this case, six).
<CFSET items=ListLen(query_string, "/")>
<CFLOOP FROM="1" TO="#items#" STEP="2" INDEX="i">
<!--- Save this URL parameter --->
<CFPARAM NAME="URL.#ListGetAt(query_string, i, "/")#"
DEFAULT="#ListGetAt(query_string, i+1, "/")#">
<!--- How many items in "query_string? --->
<!--- Loop through list, pair of items at a time --->
Next, a <CFLOOP> is used to loop through the items, from one to however many there are, stepping two at a time (because every two is a set). Within the loop <CFPARAM> is used to create the variables. The value passed to NAME is the first of the pair (i) and the value passed to DEFAULT is the second (i+1). For the first set of values (/dept/10) the <CFPARAM> would be:
<CFPARAM NAME="URL.dept" DEFAULT="10">
There you have it. By the time the </CFLOOP> is reached, a set of URL parameters would have been created. This code could be placed once at the top of a page, and all other code could refer to URL parameters without knowing they were actually faked. In addition, as <CFPARAM> was used (instead of <CFSET>), you wouldn't overwrite variables if URL parameters actually did exist.
Clean, safe, and spider friendly, too.
Custom Tags to the Rescue
Of course, you wouldn't want all that code in every page - this type of processing is ideally suited for custom tags - thus my <CF_FakeURL> tag. Listing 1 shows the complete code.
Let's take a quick look at the tag. It starts with comments and a description, as all code should. Then a <CFPARAM> is used to define the default delimiter (a slash). Next the virtual query string is extracted using two <CFSET> tags (as we did earlier).
For this code to work there must always be an even number of elements (there are two for every parameter, one for name and one for value). The next <CFIF> statement uses the MOD operator to make sure the number of items in the list is a multiple of two - if this is not the case the list is not processed.
Within the <CFLOOP> the two values (current and next) are extracted, and then <CFPARAM> creates the URL variable.
Now all you need to do is call <CF_FakeURL> in your page and any faked URL parameters will be available for you to use. When you create your URLs just change this:
Simple as that.
Arachnophobia is the irrational fear of spiders - a fear that many ColdFusion developers seem to be suffering from. As you can see, with a little ingenuity and some good old CFML, there's a very workable solution to help you overcome this fear. Enjoy!
|Daniel Leroux 12/16/03 02:09:53 PM EST|
|Vikas Tailor 07/26/03 04:18:00 PM EDT|
I used this great strategy and all of a sudden it stopped working! Could it be possible that the web server administrator no longer allows this? For example, go to http://www.linkexchanged.com/directory.cfm/Arts
This use to work, but no longer does.
|Laura Schneiderman 04/23/03 11:01:00 AM EDT|
We've noticed that Norton firewall scrambles CGI variables. This CF_FakeURL tag relies on CGI variables. Has anyone else mentioned this problem?
|Stephen Cassady 07/08/02 11:39:00 AM EDT|
MX has many undocumented changes for 5.0 to MX.
One includes the new wal URLs are examined. I have a similiar script I picked up from the developers excange which allowed me to do http://www.site.com/page.cfm/var1.value1/var2.value2/var3.value3
And it also doesn't work now. Massive blow to the way everyting works, and, ironically, to the now indexed - but non-functional pages in all the search engines.
|Seth Aaronsen 06/30/02 06:30:00 PM EDT|
We have successfully used CF_FakeURL on CF 5.0
We are having problems getting CF_FakeURL to work with CFMX standalone. This has been documented in the latest allaire message boards with the similar tag from Fusebox -- CF_FormUrl2AttributesSearch. CFMX will not process the strings http://www.mysite.com/page.cfm/variableid/value -- Please try this and you'll see.
We have tried to tweak the settings in the C:\CFusionMX\wwwroot\WEB-INF\web.xml doc by adding,
but this does not work for us.
Do you have some suggestions?
|Jill Crawford 02/22/02 09:51:00 AM EST|
Thank you for the continued information that proves to be invaluable.
I have created URLs like http://www.yoursite.com/index.cfm/productID=item similiar to what you have suggested but and have lost the query string (everything after .cfm) in the log files for IIS since doing.
Apparently ColdFusion no longer believes these are URL parameters and does not return them to IIS to write as the cs-uri-query. Is there some way to capture the paramaters to the log file?
I would greatly appreciate any assistance you can give or point me towards.
Connected devices and the industrial internet are growing exponentially every year with Cisco expecting 50 billion devices to be in operation by 2020. In this period of growth, location-based insights are becoming invaluable to many businesses as they adopt new connected technologies. Knowing when and where these devices connect from is critical for a number of scenarios in supply chain management, disaster management, emergency response, M2M, location marketing and more. In his session at @Th...
Dec. 3, 2016 09:30 AM EST Reads: 3,918
"Dice has been around for the last 20 years. We have been helping tech professionals find new jobs and career opportunities," explained Manish Dixit, VP of Product and Engineering at Dice, in this SYS-CON.tv interview at 19th Cloud Expo, held November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA.
Dec. 3, 2016 09:30 AM EST Reads: 821
What happens when the different parts of a vehicle become smarter than the vehicle itself? As we move toward the era of smart everything, hundreds of entities in a vehicle that communicate with each other, the vehicle and external systems create a need for identity orchestration so that all entities work as a conglomerate. Much like an orchestra without a conductor, without the ability to secure, control, and connect the link between a vehicle’s head unit, devices, and systems and to manage the ...
Dec. 3, 2016 09:00 AM EST Reads: 443
"We're a cybersecurity firm that specializes in engineering security solutions both at the software and hardware level. Security cannot be an after-the-fact afterthought, which is what it's become," stated Richard Blech, Chief Executive Officer at Secure Channels, in this SYS-CON.tv interview at @ThingsExpo, held November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA.
Dec. 3, 2016 08:30 AM EST Reads: 483
In addition to all the benefits, IoT is also bringing new kind of customer experience challenges - cars that unlock themselves, thermostats turning houses into saunas and baby video monitors broadcasting over the internet. This list can only increase because while IoT services should be intuitive and simple to use, the delivery ecosystem is a myriad of potential problems as IoT explodes complexity. So finding a performance issue is like finding the proverbial needle in the haystack.
Dec. 3, 2016 06:30 AM EST Reads: 6,003
In his keynote at 18th Cloud Expo, Andrew Keys, Co-Founder of ConsenSys Enterprise, provided an overview of the evolution of the Internet and the Database and the future of their combination – the Blockchain. Andrew Keys is Co-Founder of ConsenSys Enterprise. He comes to ConsenSys Enterprise with capital markets, technology and entrepreneurial experience. Previously, he worked for UBS investment bank in equities analysis. Later, he was responsible for the creation and distribution of life sett...
Dec. 3, 2016 05:45 AM EST Reads: 6,920
The WebRTC Summit New York, to be held June 6-8, 2017, at the Javits Center in New York City, NY, announces that its Call for Papers is now open. Topics include all aspects of improving IT delivery by eliminating waste through automated business models leveraging cloud technologies. WebRTC Summit is co-located with 20th International Cloud Expo and @ThingsExpo. WebRTC is the future of browser-to-browser communications, and continues to make inroads into the traditional, difficult, plug-in web ...
Dec. 3, 2016 05:30 AM EST Reads: 1,200
20th Cloud Expo, taking place June 6-8, 2017, at the Javits Center in New York City, NY, will feature technical sessions from a rock star conference faculty and the leading industry players in the world. Cloud computing is now being embraced by a majority of enterprises of all sizes. Yesterday's debate about public vs. private has transformed into the reality of hybrid cloud: a recent survey shows that 74% of enterprises have a hybrid cloud strategy.
Dec. 3, 2016 04:15 AM EST Reads: 1,718
Internet-of-Things discussions can end up either going down the consumer gadget rabbit hole or focused on the sort of data logging that industrial manufacturers have been doing forever. However, in fact, companies today are already using IoT data both to optimize their operational technology and to improve the experience of customer interactions in novel ways. In his session at @ThingsExpo, Gordon Haff, Red Hat Technology Evangelist, will share examples from a wide range of industries – includin...
Dec. 3, 2016 02:30 AM EST Reads: 1,523
WebRTC is the future of browser-to-browser communications, and continues to make inroads into the traditional, difficult, plug-in web communications world. The 6th WebRTC Summit continues our tradition of delivering the latest and greatest presentations within the world of WebRTC. Topics include voice calling, video chat, P2P file sharing, and use cases that have already leveraged the power and convenience of WebRTC.
Dec. 3, 2016 02:15 AM EST Reads: 1,501
"We build IoT infrastructure products - when you have to integrate different devices, different systems and cloud you have to build an application to do that but we eliminate the need to build an application. Our products can integrate any device, any system, any cloud regardless of protocol," explained Peter Jung, Chief Product Officer at Pulzze Systems, in this SYS-CON.tv interview at @ThingsExpo, held November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA.
Dec. 3, 2016 01:45 AM EST Reads: 788
Data is the fuel that drives the machine learning algorithmic engines and ultimately provides the business value. In his session at 20th Cloud Expo, Ed Featherston, director/senior enterprise architect at Collaborative Consulting, will discuss the key considerations around quality, volume, timeliness, and pedigree that must be dealt with in order to properly fuel that engine.
Dec. 3, 2016 12:30 AM EST Reads: 1,524
"Once customers get a year into their IoT deployments, they start to realize that they may have been shortsighted in the ways they built out their deployment and the key thing I see a lot of people looking at is - how can I take equipment data, pull it back in an IoT solution and show it in a dashboard," stated Dave McCarthy, Director of Products at Bsquare Corporation, in this SYS-CON.tv interview at @ThingsExpo, held November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA.
Dec. 2, 2016 11:15 PM EST Reads: 916
IoT is rapidly changing the way enterprises are using data to improve business decision-making. In order to derive business value, organizations must unlock insights from the data gathered and then act on these. In their session at @ThingsExpo, Eric Hoffman, Vice President at EastBanc Technologies, and Peter Shashkin, Head of Development Department at EastBanc Technologies, discussed how one organization leveraged IoT, cloud technology and data analysis to improve customer experiences and effici...
Dec. 2, 2016 08:30 PM EST Reads: 4,984
Fact is, enterprises have significant legacy voice infrastructure that’s costly to replace with pure IP solutions. How can we bring this analog infrastructure into our shiny new cloud applications? There are proven methods to bind both legacy voice applications and traditional PSTN audio into cloud-based applications and services at a carrier scale. Some of the most successful implementations leverage WebRTC, WebSockets, SIP and other open source technologies. In his session at @ThingsExpo, Da...
Dec. 2, 2016 08:15 PM EST Reads: 1,574
"IoT is going to be a huge industry with a lot of value for end users, for industries, for consumers, for manufacturers. How can we use cloud to effectively manage IoT applications," stated Ian Khan, Innovation & Marketing Manager at Solgeniakhela, in this SYS-CON.tv interview at @ThingsExpo, held November 3-5, 2015, at the Santa Clara Convention Center in Santa Clara, CA.
Dec. 2, 2016 06:45 PM EST Reads: 4,001
As data explodes in quantity, importance and from new sources, the need for managing and protecting data residing across physical, virtual, and cloud environments grow with it. Managing data includes protecting it, indexing and classifying it for true, long-term management, compliance and E-Discovery. Commvault can ensure this with a single pane of glass solution – whether in a private cloud, a Service Provider delivered public cloud or a hybrid cloud environment – across the heterogeneous enter...
Dec. 2, 2016 06:30 PM EST Reads: 1,489
The cloud promises new levels of agility and cost-savings for Big Data, data warehousing and analytics. But it’s challenging to understand all the options – from IaaS and PaaS to newer services like HaaS (Hadoop as a Service) and BDaaS (Big Data as a Service). In her session at @BigDataExpo at @ThingsExpo, Hannah Smalltree, a director at Cazena, provided an educational overview of emerging “as-a-service” options for Big Data in the cloud. This is critical background for IT and data professionals...
Dec. 2, 2016 05:00 PM EST Reads: 4,089
Today we can collect lots and lots of performance data. We build beautiful dashboards and even have fancy query languages to access and transform the data. Still performance data is a secret language only a couple of people understand. The more business becomes digital the more stakeholders are interested in this data including how it relates to business. Some of these people have never used a monitoring tool before. They have a question on their mind like “How is my application doing” but no id...
Dec. 2, 2016 04:45 PM EST Reads: 2,119
@GonzalezCarmen has been ranked the Number One Influencer and @ThingsExpo has been named the Number One Brand in the “M2M 2016: Top 100 Influencers and Brands” by Onalytica. Onalytica analyzed tweets over the last 6 months mentioning the keywords M2M OR “Machine to Machine.” They then identified the top 100 most influential brands and individuals leading the discussion on Twitter.
Dec. 2, 2016 04:45 PM EST Reads: 1,985