Click here to close now.



Welcome!

You will be redirected in 30 seconds or close now.

ColdFusion Authors: Yakov Fain, Maureen O'Gara, Nancy Y. Nee, Tad Anderson, Daniel Kaar

Related Topics: ColdFusion

ColdFusion: Article

A Cure for Arachnophobia

A Cure for Arachnophobia

Barely a week goes by without someone asking me about ColdFusion and search engine-friendly URLs. This is one of those topics that ColdFusion developers have been discussing for a long time - I first started a thread on this subject on the Allaire Developer's Forum close to five years ago. As this topic keeps coming up (and because three of you e-mailed me to ask about it this morning), I decided to scrap the column I was writing in favor of an explanation of all this once and for all.

Search Engines and ColdFusion
Here's the problem. Many search engines index sites by spidering them - starting at a known point, indexing that retrieved page, parsing it for embedded URLs, and then retrieving and indexing them too. They keep doing this until they have retrieved and indexed all linked pages in a site.

At least that's how it's supposed to work. Where things get complicated is when dynamic pages are indexed, or rather, when they're not. Some spiders won't index URLs that are dynamic, the theory being that since the content changes all the time there's no point in indexing it - after all, it may not even be the same content on subsequent accesses. So this link would be followed and indexed:

<A HREF="catalog.cfm">Catalog</A>

but this one may not:

<A HREF="catalog.cfm?dept=10">Catalog</A>

You can see the problem: ColdFusion developers (as well as ASP, PHP, Perl, Python, and JSP developers, among others) build dynamic pages. That's why we use ColdFusion - if all content was static we'd use plain old HTML without any server-side processing at all. We use ColdFusion because we want and need dynamic content - content that spiders may choose to ignore. If search engine indexing is a requirement for you, that can pose a real problem.

The Basic Solution
The solution to this problem is actually quite simple. What is it that tips off the spider, telling it that the page is dynamic? It's the query string portion of the URL, starting with the question mark. That question mark separates the URL (the file to be retrieved or the script to be executed) from any passed parameters (usually in name=value pairs).

The solution is to simply not have a query string - no question mark, no values passed after it, and no name=value pairs. Simple, eh? Yep, until you actually need passed parameters. Then what?

Well, look at this URL:

http://host/path/catalog.cfm/10

What happens when this is processed? catalog.cfm is the file to be executed and anything after the file name is ignored altogether. Even though the Web server and ColdFusion ignore it, that doesn't mean you have to. Using CGI variables (like PATH_INFO) you can access the complete URL - even parts of it that may be ignored by other processes or applications. The following snippet will display the complete path (minus any host and protocol information, although that's available in another CGI variable if needed).

<CFOUTPUT>#CGI.PATH_INFO#</CFOUTPUT>

Using the above URL this will display:

/path/catalog.cfm/10

We now know how to get the passed information. Next you need to remove the path and script name. You could do this by searching for the CFM file, but there's an easier way using yet another CGI variable, this time SCRIPT_NAME, which contains the name of the script being executed (here /path/catalog.cfm):

<CFSET query_string_length=Len(CGI.PATH_INFO)-Len(CGI.SCRIPT_NAME)>
<CFSET query_string=Right(CGI.PATH_INFO, query_string_length)>

The first <CFSET> gets the length of PATH_INFO and subtracts the length of SCRIPT_NAME from it, giving us the length of the section at the end (the data we want). The second <CFSET> uses Right() to extract the desired data, saving it in a variable named query_string. Displaying query_string would result in this output:

/10

That's how you get the data off the end of the URL.

Note: Depending on the Web server and OS being used you may find different CGI variables or different values in them. You can use <CFDUMP VAR="#CGI#"> to see all the CGI variables available to you (and you'll be able to adapt the code here accordingly).

Making It All Work
Now that the basic concept is clear, let's take this one step further. How can you pass multiple parameters? Well, you could pass three parameters, each separated by a /:

http://host/path/catalog.cfm/10/0/B12

But you would have no way of knowing which was which, what their names should be, and they'd always have to be in order (something you typically don't worry about when working with URL parameters).

While the technique is sound, the implementation can be improved by simulating name=value pairs. Look at this example:

http://host/path/catalog.cfm/dept/10/user/0/item/B12

There are now six values passed to the URL; each set of two is a name and value pair. Here it's dept=10, user=0, and item=B12. The advantage of this enhancement is that the variable name is known, the order of pairs is irrelevant, and the URL is easy to construct.

All that is required now is a simple way to extract these values and turn them into real URL values (after all, your application code shouldn't have to worry about manipulating this data; as far as it's concerned these are URL parameters plain and simple). This is where ColdFusion's list functions come in handy:

<!--- Extract "query_string" from full path --->
<CFSET query_string_length=Len(CGI.PATH_INFO)-Len(CGI.SCRIPT_NAME)>
<CFSET query_string=Right(CGI.PATH_
INFO, query_string_length)>

<!--- How many items in "query_string? --->
<CFSET items=ListLen(query_string, "/")>

<!--- Loop through list, pair of items at a time --->
<CFLOOP FROM="1" TO="#items#" STEP="2" INDEX="i">
<!--- Save this URL parameter --->
<CFPARAM NAME="URL.#ListGetAt(query_string, i, "/")#"
DEFAULT="#ListGetAt(query_string, i+1, "/")#">

</CFLOOP>

The first two <CFSET> statements simply extract the virtual query_string (as we saw earlier). Next we need to know how many items query_string contains. All ColdFusion list functions take an optional delimiter as a parameter, and so ListLen(query_string, "/") returns the number of items delimited by / (in this case, six).

Next, a <CFLOOP> is used to loop through the items, from one to however many there are, stepping two at a time (because every two is a set). Within the loop <CFPARAM> is used to create the variables. The value passed to NAME is the first of the pair (i) and the value passed to DEFAULT is the second (i+1). For the first set of values (/dept/10) the <CFPARAM> would be:

<CFPARAM NAME="URL.dept" DEFAULT="10">

There you have it. By the time the </CFLOOP> is reached, a set of URL parameters would have been created. This code could be placed once at the top of a page, and all other code could refer to URL parameters without knowing they were actually faked. In addition, as <CFPARAM> was used (instead of <CFSET>), you wouldn't overwrite variables if URL parameters actually did exist.

Clean, safe, and spider friendly, too.

Custom Tags to the Rescue
Of course, you wouldn't want all that code in every page - this type of processing is ideally suited for custom tags - thus my <CF_FakeURL> tag. Listing 1 shows the complete code.

Let's take a quick look at the tag. It starts with comments and a description, as all code should. Then a <CFPARAM> is used to define the default delimiter (a slash). Next the virtual query string is extracted using two <CFSET> tags (as we did earlier).

For this code to work there must always be an even number of elements (there are two for every parameter, one for name and one for value). The next <CFIF> statement uses the MOD operator to make sure the number of items in the list is a multiple of two - if this is not the case the list is not processed.

Within the <CFLOOP> the two values (current and next) are extracted, and then <CFPARAM> creates the URL variable.

Now all you need to do is call <CF_FakeURL> in your page and any faked URL parameters will be available for you to use. When you create your URLs just change this:

file.cfm?LName=Ben&FName=Forta

to this:

file.cfm/LName/Ben/FName/Forta

Simple as that.

Summary
Arachnophobia is the irrational fear of spiders - a fear that many ColdFusion developers seem to be suffering from. As you can see, with a little ingenuity and some good old CFML, there's a very workable solution to help you overcome this fear. Enjoy!

More Stories By Ben Forta

Ben Forta is Adobe's Senior Technical Evangelist. In that capacity he spends a considerable amount of time talking and writing about Adobe products (with an emphasis on ColdFusion and Flex), and providing feedback to help shape the future direction of the products. By the way, if you are not yet a ColdFusion user, you should be. It is an incredible product, and is truly deserving of all the praise it has been receiving. In a prior life he was a ColdFusion customer (he wrote one of the first large high visibility web sites using the product) and was so impressed he ended up working for the company that created it (Allaire). Ben is also the author of books on ColdFusion, SQL, Windows 2000, JSP, WAP, Regular Expressions, and more. Before joining Adobe (well, Allaire actually, and then Macromedia and Allaire merged, and then Adobe bought Macromedia) he helped found a company called Car.com which provides automotive services (buy a car, sell a car, etc) over the Web. Car.com (including Stoneage) is one of the largest automotive web sites out there, was written entirely in ColdFusion, and is now owned by Auto-By-Tel.

Comments (6) View Comments

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


Most Recent Comments
Daniel Leroux 12/16/03 02:09:53 PM EST
Vikas Tailor 07/26/03 04:18:00 PM EDT

Hey all,

I used this great strategy and all of a sudden it stopped working! Could it be possible that the web server administrator no longer allows this? For example, go to http://www.linkexchanged.com/directory.cfm/Arts

This use to work, but no longer does.

Thoughts?

Laura Schneiderman 04/23/03 11:01:00 AM EDT

We've noticed that Norton firewall scrambles CGI variables. This CF_FakeURL tag relies on CGI variables. Has anyone else mentioned this problem?

Stephen Cassady 07/08/02 11:39:00 AM EDT

MX has many undocumented changes for 5.0 to MX.

One includes the new wal URLs are examined. I have a similiar script I picked up from the developers excange which allowed me to do http://www.site.com/page.cfm/var1.value1/var2.value2/var3.value3

And it also doesn't work now. Massive blow to the way everyting works, and, ironically, to the now indexed - but non-functional pages in all the search engines.

Arrgh. And there are other undocumented changes too (part of the CFIDE directory beeing in the site map - something we removed for security issues - for auto JavaScript from things like CFFrom, and a CFFILE problem).

Stephen Cassady
[email protected]
http://www.lopedia.com

Seth Aaronsen 06/30/02 06:30:00 PM EDT

We have successfully used CF_FakeURL on CF 5.0

We are having problems getting CF_FakeURL to work with CFMX standalone. This has been documented in the latest allaire message boards with the similar tag from Fusebox -- CF_FormUrl2AttributesSearch. CFMX will not process the strings http://www.mysite.com/page.cfm/variableid/value -- Please try this and you'll see.

We have tried to tweak the settings in the C:\CFusionMX\wwwroot\WEB-INF\web.xml doc by adding,

CfmServlet
*

but this does not work for us.

Do you have some suggestions?

Jill Crawford 02/22/02 09:51:00 AM EST

Thank you for the continued information that proves to be invaluable.

I have created URLs like http://www.yoursite.com/index.cfm/productID=item similiar to what you have suggested but and have lost the query string (everything after .cfm) in the log files for IIS since doing.

Apparently ColdFusion no longer believes these are URL parameters and does not return them to IIS to write as the cs-uri-query. Is there some way to capture the paramaters to the log file?

I would greatly appreciate any assistance you can give or point me towards.

@ThingsExpo Stories
SYS-CON Events announced today that Commvault, a global leader in enterprise data protection and information management, has been named “Bronze Sponsor” of SYS-CON's 18th International Cloud Expo, which will take place on June 7–9, 2016, at the Javits Center in New York City, NY, and the 19th International Cloud Expo, which will take place on November 1–3, 2016, at the Santa Clara Convention Center in Santa Clara, CA. Commvault is a leading provider of data protection and information management...
The cloud promises new levels of agility and cost-savings for Big Data, data warehousing and analytics. But it’s challenging to understand all the options – from IaaS and PaaS to newer services like HaaS (Hadoop as a Service) and BDaaS (Big Data as a Service). In her session at @BigDataExpo at @ThingsExpo, Hannah Smalltree, a director at Cazena, will provide an educational overview of emerging “as-a-service” options for Big Data in the cloud. This is critical background for IT and data profes...
SYS-CON Events announced today that VAI, a leading ERP software provider, will exhibit at SYS-CON's 18th International Cloud Expo®, which will take place on June 7-9, 2016, at the Javits Center in New York City, NY. VAI (Vormittag Associates, Inc.) is a leading independent mid-market ERP software developer renowned for its flexible solutions and ability to automate critical business functions for the distribution, manufacturing, specialty retail and service sectors. An IBM Premier Business Part...
SYS-CON Events announced today that Alert Logic, Inc., the leading provider of Security-as-a-Service solutions for the cloud, will exhibit at SYS-CON's 18th International Cloud Expo®, which will take place on June 7-9, 2016, at the Javits Center in New York City, NY. Alert Logic, Inc., provides Security-as-a-Service for on-premises, cloud, and hybrid infrastructures, delivering deep security insight and continuous protection for customers at a lower cost than traditional security solutions. Ful...
Fortunately, meaningful and tangible business cases for IoT are plentiful in a broad array of industries and vertical markets. These range from simple warranty cost reduction for capital intensive assets, to minimizing downtime for vital business tools, to creating feedback loops improving product design, to improving and enhancing enterprise customer experiences. All of these business cases, which will be briefly explored in this session, hinge on cost effectively extracting relevant data from ...
With the Apple Watch making its way onto wrists all over the world, it’s only a matter of time before it becomes a staple in the workplace. In fact, Forrester reported that 68 percent of technology and business decision-makers characterize wearables as a top priority for 2015. Recognizing their business value early on, FinancialForce.com was the first to bring ERP to wearables, helping streamline communication across front and back office functions. In his session at @ThingsExpo, Kevin Roberts...
SYS-CON Events announced today that Interoute, owner-operator of one of Europe's largest networks and a global cloud services platform, has been named “Bronze Sponsor” of SYS-CON's 18th Cloud Expo, which will take place on June 7-9, 2015 at the Javits Center in New York, New York. Interoute is the owner-operator of one of Europe's largest networks and a global cloud services platform which encompasses 12 data centers, 14 virtual data centers and 31 colocation centers, with connections to 195 ad...
With an estimated 50 billion devices connected to the Internet by 2020, several industries will begin to expand their capabilities for retaining end point data at the edge to better utilize the range of data types and sheer volume of M2M data generated by the Internet of Things. In his session at @ThingsExpo, Don DeLoach, CEO and President of Infobright, will discuss the infrastructures businesses will need to implement to handle this explosion of data by providing specific use cases for filte...
As enterprises work to take advantage of Big Data technologies, they frequently become distracted by product-level decisions. In most new Big Data builds this approach is completely counter-productive: it presupposes tools that may not be a fit for development teams, forces IT to take on the burden of evaluating and maintaining unfamiliar technology, and represents a major up-front expense. In his session at @BigDataExpo at @ThingsExpo, Andrew Warfield, CTO and Co-Founder of Coho Data, will dis...
SYS-CON Events announced today that Fusion, a leading provider of cloud services, will exhibit at SYS-CON's 18th International Cloud Expo®, which will take place on June 7-9, 2016, at the Javits Center in New York City, NY. Fusion, a leading provider of integrated cloud solutions to small, medium and large businesses, is the industry's single source for the cloud. Fusion's advanced, proprietary cloud service platform enables the integration of leading edge solutions in the cloud, including clou...
Most people haven’t heard the word, “gamification,” even though they probably, and perhaps unwittingly, participate in it every day. Gamification is “the process of adding games or game-like elements to something (as a task) so as to encourage participation.” Further, gamification is about bringing game mechanics – rules, constructs, processes, and methods – into the real world in an effort to engage people. In his session at @ThingsExpo, Robert Endo, owner and engagement manager of Intrepid D...
Eighty percent of a data scientist’s time is spent gathering and cleaning up data, and 80% of all data is unstructured and almost never analyzed. Cognitive computing, in combination with Big Data, is changing the equation by creating data reservoirs and using natural language processing to enable analysis of unstructured data sources. This is impacting every aspect of the analytics profession from how data is mined (and by whom) to how it is delivered. This is not some futuristic vision: it's ha...
WebRTC has had a real tough three or four years, and so have those working with it. Only a few short years ago, the development world were excited about WebRTC and proclaiming how awesome it was. You might have played with the technology a couple of years ago, only to find the extra infrastructure requirements were painful to implement and poorly documented. This probably left a bitter taste in your mouth, especially when things went wrong.
Learn how IoT, cloud, social networks and last but not least, humans, can be integrated into a seamless integration of cooperative organisms both cybernetic and biological. This has been enabled by recent advances in IoT device capabilities, messaging frameworks, presence and collaboration services, where devices can share information and make independent and human assisted decisions based upon social status from other entities. In his session at @ThingsExpo, Michael Heydt, founder of Seamless...
The IoT's basic concept of collecting data from as many sources possible to drive better decision making, create process innovation and realize additional revenue has been in use at large enterprises with deep pockets for decades. So what has changed? In his session at @ThingsExpo, Prasanna Sivaramakrishnan, Solutions Architect at Red Hat, discussed the impact commodity hardware, ubiquitous connectivity, and innovations in open source software are having on the connected universe of people, thi...
WebRTC: together these advances have created a perfect storm of technologies that are disrupting and transforming classic communications models and ecosystems. In his session at WebRTC Summit, Cary Bran, VP of Innovation and New Ventures at Plantronics and PLT Labs, provided an overview of this technological shift, including associated business and consumer communications impacts, and opportunities it may enable, complement or entirely transform.
For manufacturers, the Internet of Things (IoT) represents a jumping-off point for innovation, jobs, and revenue creation. But to adequately seize the opportunity, manufacturers must design devices that are interconnected, can continually sense their environment and process huge amounts of data. As a first step, manufacturers must embrace a new product development ecosystem in order to support these products.
There are so many tools and techniques for data analytics that even for a data scientist the choices, possible systems, and even the types of data can be daunting. In his session at @ThingsExpo, Chris Harrold, Global CTO for Big Data Solutions for EMC Corporation, showed how to perform a simple, but meaningful analysis of social sentiment data using freely available tools that take only minutes to download and install. Participants received the download information, scripts, and complete end-t...
Manufacturing connected IoT versions of traditional products requires more than multiple deep technology skills. It also requires a shift in mindset, to realize that connected, sensor-enabled “things” act more like services than what we usually think of as products. In his session at @ThingsExpo, David Friedman, CEO and co-founder of Ayla Networks, discussed how when sensors start generating detailed real-world data about products and how they’re being used, smart manufacturers can use the dat...
When it comes to IoT in the enterprise, namely the commercial building and hospitality markets, a benefit not getting the attention it deserves is energy efficiency, and IoT’s direct impact on a cleaner, greener environment when installed in smart buildings. Until now clean technology was offered piecemeal and led with point solutions that require significant systems integration to orchestrate and deploy. There didn't exist a 'top down' approach that can manage and monitor the way a Smart Buildi...