Welcome!

You will be redirected in 30 seconds or close now.

ColdFusion Authors: Yakov Fain, Jeremy Geelan, Maureen O'Gara, Nancy Y. Nee, Tad Anderson

Related Topics: ColdFusion

ColdFusion: Article

COSMOS: Managing the ColdFusion Experience

COSMOS: Managing the ColdFusion Experience

Far too often we listen to the naysayers who tell us that something can't be done and give poorly founded reasons as to why our troubles persist. The ColdFusion Application Server is no exception to their folly. If you ask people for the drawbacks of ColdFusion, most will reply "speed" or "stability." Let me be the first to tell you that it does not have to be that way.

The system described in this article was built to change the way we think about our code and applications. COSMOS was designed to change our perceptions of the ColdFusion Application Server and to enhance the ColdFusion experience.

If you have ever looked in the /cfusion/log/ directory you've probably seen one or more of the many ColdFusion-generated error/information logs. These text files can easily grow to hundreds of MB and contain the best indicators of "what happened." As with any other service or application, a regular review of system logs should be a part of normal administration. Unfortunately, because of their large size and the fact that the data is segmented into so many logs, it's difficult to get a complete picture of performance, problems, and failure.

Developers who work on a dedicated server can use the ColdFusion Administrator to view these logs. This can be accomplished by clicking on "Log Files" and then downloading the entire log via a browser. Unfortunately, this is usually not possible given the size of most logs and the remote connection speed.

For shared developers, the critical information is unavailable due to the nature of the shared environment and security. In most cases, developers know only what a site user tells them or what they trap using CFTRY/CFCATCH and CFERROR. Even with these mechanisms in place, the larger picture is unavailable and the majority of performance issues go unnoticed and unattended.

The above issues hinder administrators and developers alike. The result is:

  • No true time or site correlation for ColdFusion Application Server events.
  • Time is wasted attempting to data mine text.
  • Site administrators, developers, and business owners don't know there is a problem.
  • A negative stigma is created based on a lack of timely and organized information.

    To be successful, a solution must have several characteristics:

  • Run autonomously, centrally, and constantly
  • Contain error lookup with "clean code" examples
  • Return all logs and bounced e-mails
  • Bring symmetry to the generated data through trending, aggregation, and normalization
  • Be fast without affecting performance on the managed server

    The solution is COSMOS. Written mainly with Cold-Fusion, it's an integration of ASP, DOS, Perl, ADSI, and Call-XML. It's a remote management platform that leverages the file system, registry, metabase, service controls, and performance counters. Currently, COSMOS contains over 16-million server events aggregated into an MS-SQL database. Captured within a maximum of 40 seconds, these events include all of the following:

    • Application errors
    • CF Application Server stop/starts
    • Hung threads
    • Long-running templates
    • Missing templates
    • Scheduled task results
    • Undeliverable e-mails
    • Mail sent
    How does this affect you? By returning timely, accurate, and relevant information, a developer can see immediately where improvement is needed. At Hostcentric, it's now possible to access several views, graphs, and aggregations that provide a new perspective on the ColdFusion experience.

    There are over 20 COSMOS reports available to a dedicated client, most of which are also available for shared customers. The following is a list of some reports with a brief description of how they impact the development and maintenance cycle.

    Information Listings
    There are several listings available, each with similar characteristics. A listing allows the user to select the maximum number of records to view per screen and how far back to examine data. It also allows the user to progress backward from that point to review previous messages. The majority of listings allow filtering to a single IIS root. They also provide direct access to the complete original error and a corrected code context lookup.

    General Application Error Listing
    Application errors are the best view into the progress and developmental completeness of a site (see Figure 1). A well-coded site generates no application errors. This listing provides a top-down view of the most recent application errors for all IIS roots. By clicking on the error message on the right, a popup window displays the error message as displayed to a site visitor.

    General Missing Template
    This applies to all .cfm templates requested by the Web server, but not found. In most cases, the developer doesn't even know that people are getting "404 File Not Found" messages. If a search engine indexes your site or a user bookmarks a page, a change in the site causes missed business. The solution is to use the default missing template handler in ColdFusion Administrator or to add a CFERROR TYPE="REQUEST" in your site's Application.cfm.

    Long-Running Template Listing
    This applies to the processing time for pages that take longer than expected. The determination of how long is too long is configured in the logging/settings section of ColdFusion Administrator. A typical setting is 45 seconds, though anything taking that long would most likely be canceled or ignored by the calling client. In addition, a script running for 45 seconds could help identify a performance bottleneck for the application server. By default, CF Administrator doesn't enable this counter. Over time, a development team should ratchet this value as low as possible to get the best diagnostics.

    Undeliverable CFMAIL Listing
    When ColdFusion is unable to deliver a message to the server specified in a CFMAIL script, the original template is renamed and filed in the /cfusion/ mail/undeliver/ directory. An error message is also written to the Mail.log or Error.log describing the problem that prevents proper delivery. This listing binds those two pieces of information together.

    The following popup allows an administrator to correct and resend the message from the original server. This function is indispensable for any business that relies on CFMAIL to reliably carry e-mail, and can't accept undelivered messages.

    Hung Thread Listing
    This is probably the greatest indicator of a performance and stability problem. Hung threads are ColdFusion's method of alerting us that it was unable to completely process the requested template. This is usually the result of code or database issues. CF4.x and above has an option in the Administrator to have CF "restart at x unresponsive requests."

    When the hung thread count matches the defined threshold, ColdFusion reaches a critical point and will stop/restart itself to avoid excessive downtime. Constant examination of hung threads is necessary to avoid application server failure. At the end of this article I've included three links that help to define more fully the causes of hung threads.

    Scheduled Task Listing
    Most scheduled tasks run completely unnoticed until someone realizes that a critical function has not processed in days. This listing is not much to look at but, under the hood, a huge modification and improvement has been created for the executive service.

    As always, COSMOS can determine if your task started, succeeded, or failed based on the logs. Furthermore, COSMOS will allow you to define a target string in the page HTML and record the generated content from the target URL to the database. If a scheduled task does not return the defined string, an e-mail containing the content and diagnostics can be generated at the time of failure. In addition, the actual HTTP response (CFHTTP.FILECONTENT) is zipped and written to the database.

    Aggregation and Stratification
    More commonly called a GROUPING, the next series of graphs were created to help identify the greatest problems quickly. By examining the data based on time, date, and IIS root, we can gather a greater understanding of where faults exist.

    Application Log Stratification by IIS Root
    Over a selectable time span, this graph allows you to see which sites are having the greatest incidence of errors (see Figure 2). By clicking on the blue horizontal bar on the right, you're driven back to the general application error listing but with an additional sort parameter that isolates errors created by the target root.

    Time/Error Graph
    Especially useful in determining if your day is getting better or worse, this graph breaks down the server errors by 10 minute increments over a selectable date span. This is often used to diagnose a recurring failure point over a multiple day or week period.

    Application Errors Stratified by Date
    Similar to the previous idea, this graph groups the number of errors by the date that they occurred (see Figure 3). This helps to identify programming trends and can easily indicate a "bad day" for an application. By clicking on the blue bar, your browser is taken to the application log stratification by IIS root. Clicking on the "Time Graph" button brings you to the next graph.

    Long-Running Template Aggregation by IIS Root
    Similar to the previous root aggregations, this has several prominent exceptions. Because a long-running page has a value associated with the processing time, I've included a column for the sum and average values. Using this display, it's possible to extract the templates most often run beyond acceptable limits, thus demanding the greatest processing time. This affects performance, though not necessarily a failure, and is a fantastic indicator of templates that need to be addressed before they become a stability issue.

    Hung Thread Aggregation by IIS Root
    This graph will often tell which application is responsible for killing the server. Over a selectable data span you can easily see which sites are causing CF to lose processing threads and tie up resources. The blue horizontal bar links back to the hung thread listing for a given root.

    The Big and the Bad
    With the thousands of errors, tasks, and events returned each hour, it's easy to become overwhelmed. In addition, not all errors have the same weight on the application server or urgency to a business owner. To resolve this problem a system of alerts and probes runs in the background. Operating at several intervals, the most relevant problems are quickly pulled out of the pool and matched with a type and severity. Once a probe identifies an event or error candidate, the alert has an option of paging, e-mailing, or calling (using CallXML) the administrator.

    A Final Look
    When did your application server last crash and why?

  • Event Chronology: As the first view that brings together data from multiple sources (see Figure 4), it provides a chronological view of all application errors, hung threads, long-running templates, and application server failures. The graph threads events are based on time in order to provide a trace leading up to a failure.
  • Spectral Analysis: This graph is unique because it rapidly identifies problems that would otherwise slip under the wire (see Figure 5). The three colors representing CF stops (red), starts (green), and hung threads (purple) are graphed relative to a 24-hour time line. By viewing all hung threads that led to server failures, a complete understanding of the root performance issues is garnered.

    Summary
    Tonight at 1 a.m. your database is going to run out of space and begin throwing application errors. Maybe your mail server stops relaying your order confirmations. There are a thousand permutations to a preventable and containable failure. Will your customers be the first to let you know?

    In the end, owners and developers, shared and dedicated, all have the same concerns: stability and performance. Using this system makes that realization no more than a few seconds away.

    Related Articles

  • http://allaire.com/Handlers/index.cfm?ID=8627&Method=Full
  • http://allaire.com/Handlers/index.cfm?ID=2497&Method=Full
  • http://allaire.com/Handlers/index.cfm?ID=1505&Method=Full
  • http://allaire.com/Handlers/index.cfm?ID=1540&Method=Full
  • http://support.microsoft.com/support/kb/articles/q174/4/96.asp
  • More Stories By Tim Nettleton

    Tim Nettleton is a senior engineer at Hostcentric’s Orlando division. He has worked with ColdFusion for four years and recently spoke at DevCon 2001 in Orlando, Florida.

    Comments (5) View Comments

    Share your thoughts on this story.

    Add your comment
    You must be signed in to add a comment. Sign-in | Register

    In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


    Most Recent Comments
    Craig Rosenblum 05/29/03 02:24:00 PM EDT

    I am not interested in the whole system mainly right now, the log component.

    It is a good idea, to help manage logs, and make sure we are really on top of errors.

    And it does look nice as well.

    jurgen koch 03/11/02 04:48:00 PM EST

    ok...let me retract the attitude from my immediately previous post and clarify the situation with more fact that flame.

    I contacted the articles author, Timothy Nettleton ([email protected]), who was very courteous and replied almost immediately to explain that while Cosmos will probably not be released for sale by Hostcentric, they can provide you remote access to Cosmos through its web interface at cosmos.hostcentric.net

    You will have to work out details re: how Hostcentric will obtain access to your CF log files for parsing, etc. but the good news is that Cosmos is available to the public...and the pricing is very reasonable.

    embarrased by my previous rant,
    Jurgen

    jurgen koch 03/11/02 01:19:00 PM EST

    I contacted Hostcentric (the ISP that the author of the article worked for) and the impression that I got was that
    COSMOS is only available to their colocation, etc. clients, and while they did develop the application, they have no plans to release it to the public.

    Kinda makes me wonder why CFDJ published the article at all. Just to tease CF Administrators?

    Hey CFDJ, want to pay me to write an article about the really cool CF apps I have developed, but can't/won't sell, or release code, logic, or the application to the public? Thanks for nothing.

    Greg Correll 02/15/02 12:55:00 PM EST

    Is cosmos available to the public?

    Matt McDonald 01/17/02 12:16:00 AM EST

    Is cosmos available to the public?

    @ThingsExpo Stories
    Internet of @ThingsExpo, taking place November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 19th Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The Internet of Things (IoT) is the most profound change in personal and enterprise IT since the creation of the Worldwide Web more than 20 years ago. All major researchers estimate there will be tens of billions devices - comp...
    The many IoT deployments around the world are busy integrating smart devices and sensors into their enterprise IT infrastructures. Yet all of this technology – and there are an amazing number of choices – is of no use without the software to gather, communicate, and analyze the new data flows. Without software, there is no IT. In this power panel at @ThingsExpo, moderated by Conference Chair Roger Strukhoff, panelists will look at the protocols that communicate data and the emerging data analy...
    “We're a global managed hosting provider. Our core customer set is a U.S.-based customer that is looking to go global,” explained Adam Rogers, Managing Director at ANEXIA, in this SYS-CON.tv interview at 18th Cloud Expo, held June 7-9, 2016, at the Javits Center in New York City, NY.
    According to Forrester Research, every business will become either a digital predator or digital prey by 2020. To avoid demise, organizations must rapidly create new sources of value in their end-to-end customer experiences. True digital predators also must break down information and process silos and extend digital transformation initiatives to empower employees with the digital resources needed to win, serve, and retain customers.
    Smart Cities are here to stay, but for their promise to be delivered, the data they produce must not be put in new siloes. In his session at @ThingsExpo, Mathias Herberts, Co-founder and CTO of Cityzen Data, will deep dive into best practices that will ensure a successful smart city journey.
    Why do your mobile transformations need to happen today? Mobile is the strategy that enterprise transformation centers on to drive customer engagement. In his general session at @ThingsExpo, Roger Woods, Director, Mobile Product & Strategy – Adobe Marketing Cloud, covered key IoT and mobile trends that are forcing mobile transformation, key components of a solid mobile strategy and explored how brands are effectively driving mobile change throughout the enterprise.
    DevOps at Cloud Expo, taking place Nov 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 19th Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The widespread success of cloud computing is driving the DevOps revolution in enterprise IT. Now as never before, development teams must communicate and collaborate in a dynamic, 24/7/365 environment. There is no time to wait for long dev...
    Cloud computing is being adopted in one form or another by 94% of enterprises today. Tens of billions of new devices are being connected to The Internet of Things. And Big Data is driving this bus. An exponential increase is expected in the amount of information being processed, managed, analyzed, and acted upon by enterprise IT. This amazing is not part of some distant future - it is happening today. One report shows a 650% increase in enterprise data by 2020. Other estimates are even higher....
    The Jevons Paradox suggests that when technological advances increase efficiency of a resource, it results in an overall increase in consumption. Writing on the increased use of coal as a result of technological improvements, 19th-century economist William Stanley Jevons found that these improvements led to the development of new ways to utilize coal. In his session at 19th Cloud Expo, Mark Thiele, Chief Strategy Officer for Apcera, will compare the Jevons Paradox to modern-day enterprise IT, e...
    What happens when the different parts of a vehicle become smarter than the vehicle itself? As we move toward the era of smart everything, hundreds of entities in a vehicle that communicate with each other, the vehicle and external systems create a need for identity orchestration so that all entities work as a conglomerate. Much like an orchestra without a conductor, without the ability to secure, control, and connect the link between a vehicle’s head unit, devices, and systems and to manage the ...
    In this strange new world where more and more power is drawn from business technology, companies are effectively straddling two paths on the road to innovation and transformation into digital enterprises. The first path is the heritage trail – with “legacy” technology forming the background. Here, extant technologies are transformed by core IT teams to provide more API-driven approaches. Legacy systems can restrict companies that are transitioning into digital enterprises. To truly become a lea...
    What are the new priorities for the connected business? First: businesses need to think differently about the types of connections they will need to make – these span well beyond the traditional app to app into more modern forms of integration including SaaS integrations, mobile integrations, APIs, device integration and Big Data integration. It’s important these are unified together vs. doing them all piecemeal. Second, these types of connections need to be simple to design, adapt and configure...
    Information technology is an industry that has always experienced change, and the dramatic change sweeping across the industry today could not be truthfully described as the first time we've seen such widespread change impacting customer investments. However, the rate of the change, and the potential outcomes from today's digital transformation has the distinct potential to separate the industry into two camps: Organizations that see the change coming, embrace it, and successful leverage it; and...
    19th Cloud Expo, taking place November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA, will feature technical sessions from a rock star conference faculty and the leading industry players in the world. Cloud computing is now being embraced by a majority of enterprises of all sizes. Yesterday's debate about public vs. private has transformed into the reality of hybrid cloud: a recent survey shows that 74% of enterprises have a hybrid cloud strategy. Meanwhile, 94% of enterpri...
    SYS-CON Events announced today that CDS Global Cloud, an Infrastructure as a Service provider, will exhibit at the 19th International Cloud Expo, which will take place on November 1–3, 2016, at the Santa Clara Convention Center in Santa Clara, CA. CDS Global Cloud is an IaaS (Infrastructure as a Service) provider specializing in solutions for e-commerce, internet gaming, online education and other internet applications. With a growing number of data centers and network points around the world, ...
    In his general session at 18th Cloud Expo, Lee Atchison, Principal Cloud Architect and Advocate at New Relic, discussed cloud as a ‘better data center’ and how it adds new capacity (faster) and improves application availability (redundancy). The cloud is a ‘Dynamic Tool for Dynamic Apps’ and resource allocation is an integral part of your application architecture, so use only the resources you need and allocate /de-allocate resources on the fly.
    Major trends and emerging technologies – from virtual reality and IoT, to Big Data and algorithms – are helping organizations innovate in the digital era. However, to create real business value, IT must think beyond the ‘what’ of digital transformation to the ‘how’ to harness emerging trends, innovation and disruption. Architecture is the key that underpins and ties all these efforts together. In the digital age, it’s important to invest in architecture, extend the enterprise footprint to the cl...
    There are several IoTs: the Industrial Internet, Consumer Wearables, Wearables and Healthcare, Supply Chains, and the movement toward Smart Grids, Cities, Regions, and Nations. There are competing communications standards every step of the way, a bewildering array of sensors and devices, and an entire world of competing data analytics platforms. To some this appears to be chaos. In this power panel at @ThingsExpo, moderated by Conference Chair Roger Strukhoff, Bradley Holt, Developer Advocate a...
    SYS-CON Events announced today that LeaseWeb USA, a cloud Infrastructure-as-a-Service (IaaS) provider, will exhibit at the 19th International Cloud Expo, which will take place on November 1–3, 2016, at the Santa Clara Convention Center in Santa Clara, CA. LeaseWeb is one of the world's largest hosting brands. The company helps customers define, develop and deploy IT infrastructure tailored to their exact business needs, by combining various kinds cloud solutions.
    A strange thing is happening along the way to the Internet of Things, namely far too many devices to work with and manage. It has become clear that we'll need much higher efficiency user experiences that can allow us to more easily and scalably work with the thousands of devices that will soon be in each of our lives. Enter the conversational interface revolution, combining bots we can literally talk with, gesture to, and even direct with our thoughts, with embedded artificial intelligence, wh...